Simultaneous interpretation method, device, system and equipment

By acquiring the pronunciation features and emotion categories of speech, the machine imitates the pronunciation features and emotional states to generate translation results, solving the problems of high cost, low efficiency and insufficient emotional expression in traditional simultaneous interpretation, and achieving efficient and authentic cross-language communication.

CN120977288APending Publication Date: 2025-11-18HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410796632.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-05-17
Filing Date
2024-06-19
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Traditional human simultaneous interpretation suffers from high costs in training and working interpreters, low translation efficiency, untimely translation, and difficulties in translating multiple languages ​​simultaneously. Furthermore, the voice output by machine simultaneous interpretation differs greatly from the speaker's voice, failing to convey the speaker's true emotions and affecting the translation effect.

Method used

By acquiring the pronunciation features and emotion categories of the speech, the machine imitates the pronunciation features and emotional states to generate a translation result that is consistent with the original speech. Optionally, it can be combined with synchronized playback of the person's video to achieve audio-visual synchronization.

Benefits of technology

It improves translation quality and user experience, making translation results more authentic and achieving the same language effect in cross-language communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977288A_ABST
    Figure CN120977288A_ABST
Patent Text Reader

Abstract

The invention provides a simultaneous interpretation method, device, system and equipment, and belongs to the field of artificial intelligence. Firstly, a character feature, an emotion category and a first text corresponding to a first voice are obtained; the character features comprise pronunciation features. Secondly, translating the first text in the first language into a second text in a second language, wherein the second language is different from the first language; and generating a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice and the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice, and the emotional state reflected by the second voice is the same as the emotional state reflected by the first voice. And finally, outputting a simultaneous interpretation result corresponding to the first voice, wherein the simultaneous interpretation result comprises the second voice. By simulating the pronunciation characteristics and the emotion types corresponding to the original voice, the translation result is output according to the pronunciation characteristics and the emotion types corresponding to the original voice, and the translation effect and the user experience are improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the Chinese Patent Application No. 202410636958.3, filed on May 17, 2024, and entitled "A Real-Time Simultaneous Interpretation Audio-Video Conference Method and Device", the entire contents of which are incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the field of artificial intelligence, and in particular to a simultaneous interpretation method, device, system and equipment. BACKGROUND

[0003] Nowadays, more and more conferences involve participants from multiple different countries. In order for all participants to understand the conference content, it is often necessary to perform simultaneous interpretation on the conference content. Traditional simultaneous interpretation is mostly performed by manual translation. However, manual translation has problems such as high cost of interpreter training and work, low translation efficiency, non-timely translation, and difficulty in multi-language alternating translation.

[0004] With the rapid development of scientific technologies such as artificial intelligence (AI), machine simultaneous interpretation has begun to be applied in multi-person multi-point conferences. Machine simultaneous interpretation is a technical direction in the field of intelligent voice interaction. Machine simultaneous interpretation has advantages such as high recognition and memory integrity of objective information, rich terminology and vocabulary library, and low-cost long-time efficient work. Using machine simultaneous interpretation to assist or even replace manual translation has become a trend. However, how to improve the translation effect of machine simultaneous interpretation is a problem to be solved at present. SUMMARY

[0005] The present application provides a simultaneous interpretation method, device, system and equipment.

[0006] In a first aspect, a simultaneous interpretation method is provided. The method is applied to a computing device. The method includes: obtaining a character feature corresponding to a first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice. The character feature includes a pronunciation feature. The first text is used to describe the voice content of the first voice. The first text is translated into a second text, the first text is in a first language, and the second text is in a second language, the second language being different from the first language. According to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, a second voice is generated. The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice, and the emotion state reflected by the second voice is the same as the emotion state reflected by the first voice. A simultaneous interpretation result corresponding to the first voice is output, and the simultaneous interpretation result includes the second voice.

[0007] The application obtains pronunciation features and emotion categories corresponding to the first speech, and in the process of translating the first speech in the first language into the second speech in the second language, the pronunciation features and emotion categories corresponding to the first speech are imitated by the machine, so that the translation result is output in the pronunciation features and emotion categories corresponding to the first speech, that is, the second speech is consistent with the pronunciation features corresponding to the first speech and the same emotion state is reflected, so that the translation result is more authentic and the translation effect is improved. In addition, the same language communication effect is realized in the cross-language communication scene, thereby improving the user experience.

[0008] Optionally, an implementation of obtaining the character features corresponding to the first speech includes: performing voiceprint recognition on the first speech to obtain target voiceprint features corresponding to the first speech. The target voiceprint features are matched with voiceprint features in a target feature set library, and the target feature set library includes a plurality of corresponding relationships between voiceprint features and character features. If the target voiceprint features are included in the target feature set library, the character features corresponding to the target voiceprint features in the target feature set library are taken as the character features corresponding to the first speech.

[0009] In the application, in the case where the target voiceprint features are included in the target feature set library, the computing device directly obtains the pronunciation features corresponding to the first speech from the target feature set library through voiceprint feature matching, and the efficiency of obtaining the pronunciation features is high.

[0010] Optionally, if the target voiceprint features are not included in the target feature set library, pronunciation features of the first speech are extracted to obtain the pronunciation features corresponding to the first speech. The target voiceprint features and the pronunciation features corresponding to the first speech are stored in the target feature set library.

[0011] In the application, in the case where the target voiceprint features are not included in the target feature set library, the computing device obtains the pronunciation features corresponding to the first speech by extracting the pronunciation features of the first speech, and realizes real-time acquisition of the pronunciation features. The corresponding relationship between the voiceprint features and the pronunciation features in the target feature set library can be obtained from real-time speech of the user.

[0012] Optionally, the above method further includes: obtaining user registration information, the user registration information including a recorded speech. Voiceprint recognition is performed on the recorded speech to obtain voiceprint features corresponding to the recorded speech, and pronunciation features of the recorded speech are extracted to obtain pronunciation features corresponding to the recorded speech. The voiceprint features corresponding to the recorded speech and the pronunciation features corresponding to the recorded speech are stored in the target feature set library.

[0013] In the application, the corresponding relationship between the voiceprint features and the pronunciation features in the target feature set library can be obtained from the recorded speech provided by the user when the user registers.

[0014] Optionally, the character feature further includes a character image, and the character image includes a face of the target character. The method further includes: performing speech driving on the character image by using the second speech to obtain a character video, and a mouth movement of the target character in the character video matches the second speech. Correspondingly, the simultaneous interpretation result further includes the character video. In this case, the user registration information further includes a recording video or a captured image, the character image is collected from the recording video or the captured image, and the target feature set library is used to store a correspondence between a plurality of sets of voiceprint features and pronunciation features and a character image.

[0015] The application outputs a character video in which a mouth movement of a character matches the translated speech, so that the speech and the character video are played simultaneously, the mouth movement of the character in the speech and the character video is consistent, and thus audio and video are played synchronously, and user experience is improved.

[0016] Optionally, an implementation of obtaining the emotion category corresponding to the first speech includes: performing emotion recognition on the first speech to obtain the emotion category corresponding to the first speech.

[0017] Optionally, an implementation of generating the second speech according to the pronunciation feature corresponding to the first speech, the emotion category corresponding to the first speech, and the second text includes: inputting the pronunciation feature corresponding to the first speech, the emotion category corresponding to the first speech, and the second text into a speech synthesis model to obtain the second speech output by the speech synthesis model.

[0018] The speech synthesis model is a machine learning model trained based on a plurality of sets of training samples. Optionally, each set of training samples includes sample pronunciation features, an emotion category label, and sample text. Alternatively, each set of training samples includes sample pronunciation features, an emotion category label, sample text, and a speech label.

[0019] Optionally, the method is applied to a conference system, and the first speech is from a first conference terminal in the conference system. An implementation of outputting the simultaneous interpretation result corresponding to the first speech includes: sending the simultaneous interpretation result to a second conference terminal in the conference system.

[0020] Optionally, the method further includes: receiving a language selection instruction sent by the second conference terminal, and the language selection instruction is used to indicate selection of the second language.

[0021] In a second aspect, a simultaneous interpretation apparatus is provided. The apparatus includes a plurality of functional modules that interact with each other to implement the method in the first aspect and each of the implementations. The plurality of functional modules can be implemented based on software, hardware, or a combination of software and hardware, and the plurality of functional modules can be combined or divided based on specific implementation.

[0022] In a third aspect, a conference system is provided, which includes a simultaneous interpretation server, a conference service platform, and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The conference service platform is connected to the simultaneous interpretation server. The plurality of conference terminals includes a first conference terminal and a second conference terminal.

[0023] The conference service platform is configured to receive a first voice sent by the first conference terminal and send the first voice to the simultaneous interpretation server. The simultaneous interpretation server is configured to obtain a character feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text. The character feature includes a pronunciation feature. The first text is used to describe a voice content of the first voice. The first text is in a first language. The second text is in a second language. The second language is different from the first language. The simultaneous interpretation server is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform. The simultaneous interpretation result includes the second voice. A voice content of the second voice is a description content of the second text. A pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice. An emotion state reflected by the second voice is the same as an emotion state reflected by the first voice. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

[0024] In a fourth aspect, another conference system is provided, which includes a simultaneous interpretation server, a conference service platform, a virtual interpreter terminal, and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The virtual interpreter terminal is connected to the conference service platform and the simultaneous interpretation server, respectively. The plurality of conference terminals includes a first conference terminal and a second conference terminal.

[0025] The conference service platform is configured to receive the first voice sent by the first conference terminal and send the first voice to a virtual interpreter terminal. The virtual interpreter terminal is configured to send the first voice to a simultaneous interpretation server. The simultaneous interpretation server is configured to obtain a character feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text. The character feature includes a pronunciation feature. The first text is used to describe the voice content of the first voice. The first text is in a first language, and the second text is in a second language. The second language is different from the first language. The simultaneous interpretation server is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the virtual interpreter terminal. The simultaneous interpretation result includes the second voice. The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice. The emotion state reflected by the second voice is the same as the emotion state reflected by the first voice. The virtual interpreter terminal is configured to send the simultaneous interpretation result to the conference service platform. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

[0026] Optionally, in combination with the third aspect or the fourth aspect, the conference system has a plurality of simultaneous interpretation engines deployed in the simultaneous interpretation server. Each of the simultaneous interpretation engines corresponds to one conference terminal in the conference system, or each of the simultaneous interpretation engines corresponds to one conference in the conference system, or each of the simultaneous interpretation engines corresponds to one language. The language is a language to be translated or a target language to be translated.

[0027] In the present application, the simultaneous interpretation engines in the simultaneous interpretation server are deployed according to the conference terminal, the number of conferences, or the language category. One simultaneous interpretation engine provides simultaneous interpretation services for one conference terminal, one conference, or one language. This solves the problems of insecurity, complex routing, vulnerability to attack, string translation, poor disaster recovery, and mutual interference of multiple requests when one simultaneous interpretation engine corresponds to multiple conference terminals, multiple conferences, or multiple languages, thereby improving the security, flexibility, and practicality of the simultaneous interpretation engine.

[0028] In a fifth aspect, another conference system is provided. The conference system includes a conference service platform and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The plurality of conference terminals include a first conference terminal and a second conference terminal.

[0029] The conference service platform is configured to receive a first voice sent by the first conference terminal, acquire a character feature corresponding to the first voice, and send the first voice and the character feature corresponding to the first voice to the second conference terminal, the character feature including a pronunciation feature. The second conference terminal is configured to acquire an emotion category corresponding to the first voice and a first text corresponding to the first voice, and translate the first text into a second text, the first text being used to describe a voice content of the first voice, the first text being in a first language, and the second text being in a second language, the second language being different from the first language. The second conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and play a simultaneous interpretation result corresponding to the first voice, the simultaneous interpretation result including the second voice, a voice content of the second voice being description content of the second text, a pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and an emotion state reflected by the second voice being the same as an emotion state reflected by the first voice.

[0030] Optionally, in combination with the fifth aspect, the conference system further includes a simultaneous interpretation server, the conference service platform is connected with the simultaneous interpretation server, and the simultaneous interpretation server stores a target feature set library, the target feature set library including a plurality of corresponding relationships between voiceprint features and character features. The conference service platform is configured to perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice, and send the target voiceprint feature to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint feature with voiceprint features in the target feature set library, and if the target feature set library includes the target voiceprint feature, send a character feature corresponding to the target voiceprint feature in the target feature set library to the conference service platform. The conference service platform is configured to take the character feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the character feature corresponding to the first voice.

[0031] In a sixth aspect, another conference system is provided, which includes a conference service platform and a plurality of conference terminals, the plurality of conference terminals being in communication through the conference service platform, and the plurality of conference terminals including a first conference terminal and a second conference terminal.

[0032] The conference service platform is configured to receive a first voice sent by a first conference terminal and send the first voice to a second conference terminal. The second conference terminal is configured to obtain a character feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text. The character feature includes a pronunciation feature. The first text is used to describe the voice content of the first voice. The first text is in a first language, and the second text is in a second language. The second language is different from the first language. The second conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and play a simultaneous interpretation result corresponding to the first voice. The simultaneous interpretation result includes the second voice. The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice. The emotion state reflected by the second voice is the same as the emotion state reflected by the first voice.

[0033] Optionally, in combination with the sixth aspect, the conference system further includes a simultaneous interpretation server, the plurality of conference terminals are respectively connected with the simultaneous interpretation server, and the simultaneous interpretation server stores a target feature set library. The target feature set library includes a plurality of corresponding relationships between voiceprint features and character features. The second conference terminal is configured to perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice, and send the target voiceprint feature to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint feature with voiceprint features in the target feature set library. If the target feature set library includes the target voiceprint feature, the simultaneous interpretation server sends a character feature corresponding to the target voiceprint feature in the target feature set library to the second conference terminal. The second conference terminal is configured to take the character feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the character feature corresponding to the first voice.

[0034] In a seventh aspect, another conference system is provided. The conference system includes a conference service platform and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The plurality of conference terminals include a first conference terminal and a second conference terminal.

[0035] The first conference terminal is configured to collect the first voice and send the first voice to the conference service platform. The conference service platform is configured to obtain a character feature corresponding to the first voice, and send the character feature corresponding to the first voice to the second conference terminal. The character feature includes a pronunciation feature. The first conference terminal is configured to obtain an emotion category corresponding to the first voice and a first text corresponding to the first voice, and translate the first text into a second text. The first text is used to describe the voice content of the first voice. The first text is in a first language, and the second text is in a second language. The second language is different from the first language. The first conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform. The simultaneous interpretation result includes the second voice. The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice. The emotion state reflected by the second voice is the same as the emotion state reflected by the first voice. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

[0036] Optionally, in combination with the seventh aspect, the conference system further includes a simultaneous interpretation server. The conference service platform is connected with the simultaneous interpretation server. The target feature set library is stored in the simultaneous interpretation server. The target feature set library includes a plurality of corresponding relationships between voiceprint features and character features. The conference service platform is configured to perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice, and send the target voiceprint feature to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint feature with the voiceprint features in the target feature set library. If the target feature set library includes the target voiceprint feature, the simultaneous interpretation server sends the character feature corresponding to the target voiceprint feature in the target feature set library to the conference service platform. The conference service platform is configured to take the character feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the character feature corresponding to the first voice.

[0037] In an eighth aspect, another conference system is provided. The conference system includes a conference service platform and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The plurality of conference terminals include a first conference terminal and a second conference terminal.

[0038] The first conference terminal is configured to collect the first voice, obtain a character feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text. The character feature includes a pronunciation feature. The first text is used to describe the voice content of the first voice. The first text is in a first language, and the second text is in a second language. The second language is different from the first language. The first conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform. The simultaneous interpretation result includes the second voice. The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice. The emotion state reflected by the second voice is the same as the emotion state reflected by the first voice. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

[0039] Optionally, in combination with the eighth aspect, the conference system further includes a simultaneous interpretation server. The first conference terminal is connected with the simultaneous interpretation server. The simultaneous interpretation server stores a target feature set library. The target feature set library includes a plurality of corresponding relationships between voiceprint features and character features. The first conference terminal is configured to perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice, and send the target voiceprint feature to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint feature with the voiceprint features in the target feature set library. If the target feature set library includes the target voiceprint feature, the simultaneous interpretation server sends a character feature corresponding to the target voiceprint feature in the target feature set library to the first conference terminal. The first conference terminal is configured to use the character feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the character feature corresponding to the first voice.

[0040] Optionally, in combination with any one of the third aspect to the eighth aspect, the first conference terminal is configured to display a conference registration interface. The conference registration interface displays a character information input option. The character information input option is used for a user to provide character information. The character information includes a recorded voice. The first conference terminal is further configured to generate user registration information in response to receiving the recorded voice through the conference registration interface. The user registration information includes the recorded voice. The recorded voice is used to determine a corresponding relationship between a group of voiceprint features and pronunciation features.

[0041] Optionally, in combination with any one of the third aspect to the eighth aspect, the second conference terminal is configured to receive a language selection instruction through a display interface. The language selection instruction is used to indicate selection of the second language.

[0042] In a ninth aspect, a computing device cluster is provided, including at least one computing device, each computing device including a processor and a memory. The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the method in the first aspect and each of the implementations thereof.

[0043] In a tenth aspect, a computer program product including instructions is provided, which, when executed by a computing device cluster, causes the computing device cluster to perform the method in the first aspect and each of the implementations thereof.

[0044] In an eleventh aspect, a computer-readable storage medium is provided, including computer program instructions, which, when executed by a computing device cluster, causes the computing device cluster to perform the method in the first aspect and each of the implementations thereof.

[0045] In a twelfth aspect, a chip is provided, including programmable logic circuit and / or program instructions, which, when the chip is executed, implements the method in the first aspect and each of the implementations thereof. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 is a flowchart of a simultaneous interpretation method provided by an embodiment of the present application;

[0047] Figure 2 is a functional module diagram of a simultaneous interpretation engine provided by an embodiment of the present application;

[0048] Figure 3 is a flowchart of another simultaneous interpretation method provided by an embodiment of the present application;

[0049] Figure 4 is a functional module diagram of another simultaneous interpretation engine provided by an embodiment of the present application;

[0050] Figure 5 is a video conference scene diagram provided by an embodiment of the present application;

[0051] Figure 6 is a display interface diagram provided by an embodiment of the present application;

[0052] Figure 7 is a structure diagram of a conference system provided by an embodiment of the present application;

[0053] Figure 8 is a structure diagram of another conference system provided by an embodiment of the present application;

[0054] Figure 9 is a structure diagram of yet another conference system provided by an embodiment of the present application;

[0055] Figure 10 is a structural schematic diagram of still another conference system provided by an embodiment of the present application;

[0056] Figure 11 is a structural schematic diagram of still another conference system provided by an embodiment of the present application;

[0057] Figure 12 is a structural schematic diagram of still another conference system provided by an embodiment of the present application;

[0058] Figure 13 is a structural schematic diagram of a simultaneous interpretation device provided by an embodiment of the present application;

[0059] Figure 14 is a structural schematic diagram of a computing device provided by an embodiment of the present application;

[0060] Figure 15 is a structural schematic diagram of a computing device cluster provided by an embodiment of the present application;

[0061] Figure 16 is a structural schematic diagram of another computing device cluster provided by an embodiment of the present application. DETAILED DESCRIPTION

[0062] In order to make the objects, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.

[0063] Machine simultaneous interpretation refers to converting the speech of a speaker into text in real time through advanced speech recognition technology, and then immediately translating the text into a target language using machine translation technology to achieve almost no-delay simultaneous interpretation. Machine simultaneous interpretation effectively solves the problems of high cost of interpreter training and work, low translation efficiency, non-timely translation, and difficulty in multi-lingual alternating translation, etc. existing in current manual translation, but there is still a lot of room for improvement. For example, the output speech of machine simultaneous interpretation is usually mechanical and stereotyped, which is obviously different from the speaker's voice, and cannot express the real emotions of the speaker when speaking, which may affect the translation effect and result in poor user experience.

[0064] Based on this, the technical scheme provided in the present application combines the pronunciation characteristics of the speaker and the emotional state when speaking in the machine simultaneous interpretation process, so that the output voice after machine simultaneous interpretation is the same as the pronunciation and emotional state when speaking of the speaker, making the translation result more authentic and improving the translation effect. In the cross-language communication scene, the effect of communication in the same language is realized, thereby improving the user experience. The technical scheme provided in the present application is applied to a computing device, and the technical scheme is specifically implemented as follows. First, the characteristics of the person corresponding to the first voice, the emotional category corresponding to the first voice, and the first text corresponding to the first voice are obtained. The characteristics of the person include pronunciation characteristics, and the first text is used to describe the voice content of the first voice. Second, the first text is translated into a second text. The first text uses a first language, and the second text uses a second language, which is different from the first language. Then, the second voice is generated according to the pronunciation characteristics corresponding to the first voice, the emotional category corresponding to the first voice, and the second text. The voice content of the second voice is the description content of the second text. The pronunciation characteristics corresponding to the second voice are consistent with the pronunciation characteristics corresponding to the first voice, and the emotional state reflected by the second voice is the same as the emotional state reflected by the first voice. Finally, the simultaneous interpretation result corresponding to the first voice is output, and the simultaneous interpretation result includes the second voice. The above first voice is the original voice to be translated, and the first text is the original text converted from the original voice. The second text is the target text translated from the original text, and the second voice is the target voice converted from the target text. The present application obtains the pronunciation characteristics and emotional category corresponding to the original voice, and in the process of translating the original voice into the target voice in a specified language, the machine imitates the pronunciation characteristics and emotional category corresponding to the original voice, so that the translation result is output in the pronunciation characteristics and emotional category corresponding to the original voice, that is, the target voice is consistent with the pronunciation characteristics corresponding to the original voice and the emotional state reflected is the same, so that the translation result is more authentic, thereby improving the translation effect and user experience.

[0065] Optionally, the pronunciation characteristics include timbre characteristics. In this way, the machine can imitate the timbre of the speaker to output the target voice, so that the listener hears the speaker speaking in the target language, thereby improving the user experience. Further, the pronunciation characteristics can also include pitch and / or speech rate.

[0066] Optionally, in some embodiments, the character feature corresponding to the first speech further includes a character image, the character image including a face of the target character. After the second speech is generated, the character image is further speech-driven using the second speech to obtain a character video. The mouth movement of the target character in the character video matches the second speech. The simultaneous interpretation result further includes the character video. The mouth movement of the target character in the character video matches the second speech means that the mouth movement of the target character is the same as the mouth movement when the speaking content of the target character is the second speech. Through outputting the character video in which the mouth movement of the character matches the translated speech, the speech and the mouth movement of the character in the character video are consistent when the speech and the character video are played, so that the audio and video are played synchronously, and the user experience is improved.

[0067] Optionally, the character image is a real person image or a digital person. The digital person refers to a digital character image (virtual character image) close to the image of a human being created by computer technology. When the character image is a real person image, for example, when the character image is a speaker image providing the first speech, the application can provide a high-fidelity audio and video of the speaker speaking in a non-native language (second language) for the user to watch, thereby improving the user experience.

[0068] The method provided by the embodiments of the application can be applied to various scenes requiring simultaneous interpretation, including but not limited to video conference scenes, online translation, remote education, online live broadcast, and the like.

[0069] The method flow of the embodiments of the application is illustrated below.

[0070] The simultaneous interpretation method provided by the embodiments of the application can output only audio, or can output audio and video. The implementation processes of the two schemes are described below by methods 100 and 300 respectively.

[0071] For example, Figure 1 is a flowchart of a simultaneous interpretation method 100 provided by the embodiments of the application. In the method 100, only audio is output. As Figure 1 shown, the method 100 includes but is not limited to the following steps 101 to 104.

[0072] In step 101, the computing device obtains a character feature corresponding to the first speech, an emotion category corresponding to the first speech, and a first text corresponding to the first speech, the character feature including a pronunciation feature.

[0073] The first text is used to describe the speech content of the first speech, i.e., the speech content of the first speech is the description content of the first text. The first speech is the original speech to be translated. The first speech and the first text are in the first language. Optionally, the pronunciation features include at least timbre features. Further, the pronunciation features also include tone and / or speech rate.

[0074] Optionally, the computing device obtains the implementation of the character feature corresponding to the first speech, including but not limited to the following steps 1011 to 1013. Further, the following steps 1014 to 1015 can also be included.

[0075] In step 1011, the computing device performs voiceprint recognition on the first speech to obtain the target voiceprint feature corresponding to the first speech.

[0076] Voiceprint, also known as speech fingerprint, is a stable biometric feature. Everyone has a unique voiceprint, similar to a fingerprint, which can be used as an identity identifier to identify the identity of a person.

[0077] In step 1012, the computing device matches the target voiceprint feature with voiceprint features in a target feature set library, which includes a plurality of voiceprint features and corresponding character features.

[0078] In this method 100, the target feature set library includes a plurality of voiceprint features and corresponding pronunciation features, and the target feature set library is a pronunciation feature set library. For example, the target feature set library stores a plurality of voiceprint features and corresponding pronunciation features in the form of key-value pairs, with voiceprint features as keys and pronunciation features as values. The target feature set library can be represented as follows: {voiceprint feature 1: pronunciation feature 1; voiceprint feature 2: pronunciation feature 2; voiceprint feature 3: pronunciation feature 3; …}.

[0079] Optionally, the correspondence between the voiceprint features and the pronunciation features in the target feature set library can be pre-acquired from a recorded speech provided by a user during user registration. The computing device obtains user registration information, which includes a recorded speech. The computing device performs voiceprint recognition on the recorded speech to obtain a voiceprint feature corresponding to the recorded speech, and performs pronunciation feature extraction on the recorded speech to obtain a pronunciation feature corresponding to the recorded speech. The computing device stores the voiceprint feature corresponding to the recorded speech and the pronunciation feature corresponding to the recorded speech in the target feature set library.

[0080] Alternatively, the corresponding relationship between the voiceprint feature and the pronunciation feature in the target feature set library can also be obtained from real-time speech of the user. For example, when the voiceprint feature of a speaker is not included in the target feature set library, the pronunciation feature of the speaker is extracted from the speech content of the speaker to obtain the pronunciation feature of the speaker, and the voiceprint feature and the pronunciation feature of the speaker are further stored correspondingly in the target feature set library. After that, when the speaker speaks again, since the voiceprint feature of the speaker is already included in the target feature set library, the computing device does not need to repeatedly extract the pronunciation feature of the speaker, thereby improving the efficiency of simultaneous interpretation.

[0081] In step 1013, if the target voiceprint feature is included in the target feature set library, the computing device takes the person feature corresponding to the target voiceprint feature in the target feature set library as the person feature corresponding to the first speech.

[0082] In the embodiments of the present application, when the target voiceprint feature is included in the target feature set library, the computing device directly obtains the pronunciation feature corresponding to the first speech from the target feature set library in a voiceprint feature matching manner, and the efficiency of obtaining the pronunciation feature is high.

[0083] In step 1014, if the target voiceprint feature is not included in the target feature set library, the computing device extracts the pronunciation feature of the first speech to obtain the pronunciation feature corresponding to the first speech.

[0084] In the embodiments of the present application, when the target voiceprint feature is not included in the target feature set library, the computing device obtains the pronunciation feature corresponding to the first speech by extracting the pronunciation feature of the first speech, and realizes real-time acquisition of the pronunciation feature.

[0085] In step 1015, the computing device correspondingly stores the target voiceprint feature and the pronunciation feature corresponding to the first speech in the target feature set library.

[0086] In the embodiments of the present application, when the target voiceprint feature is not included in the target feature set library, after obtaining the pronunciation feature corresponding to the first speech, the computing device correspondingly stores the target voiceprint feature and the pronunciation feature corresponding to the first speech in the target feature set library to obtain an updated target feature set library. In this way, when the speaker providing the first speech speaks again, the computing device can quickly match the pronunciation feature of the speaker based on the updated target feature set library, without the need to extract the pronunciation feature of the speech content of the speaker again, thereby improving the efficiency of simultaneous interpretation.

[0087] Optionally, the implementation manner in which the computing device obtains the emotion category corresponding to the first voice includes: the computing device performing emotion recognition on the first voice to obtain the emotion category corresponding to the first voice. The emotion category includes, but is not limited to, default, angry, whispering, excited, cheerful, sad, or friendly.

[0088] Optionally, the implementation manner in which the computing device obtains the first text corresponding to the first voice includes: the computing device converting the first voice into the first text by using a speech to text (STT) technology. For example, the computing device converts the voice into the text by using an automatic speech recognition (ASR) technology.

[0089] Step 102, the computing device translates the first text into a second text.

[0090] The first text is in a first language, and the second text is in a second language different from the first language. The computing device translates the first text in the first language into the second text in the second language by using a machine translation algorithm. Optionally, the machine translation algorithm used by the computing device includes, but is not limited to, a statistical machine translation (SMT) algorithm, a neural machine translation (NMT) algorithm, a pretrained language models algorithm, an attention-based models algorithm, or a reinforcement learning algorithm.

[0091] Step 103, the computing device generates a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text.

[0092] The voice content of the second voice is the description content of the second text, that is, the second text is used to describe the voice content of the second voice. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice, and the emotion state reflected by the second voice is the same as the emotion state reflected by the first voice. The computing device can convert the second text into the second voice by using a text to speech (TTS) technology combined with a voice cloning technology.

[0093] Optionally, in an implementation of the step 103, the first speech corresponding pronunciation feature, the first speech corresponding emotion category and the second text are input into a speech synthesis model to obtain the second speech output by the speech synthesis model. The speech synthesis model is a machine learning model trained based on a plurality of training samples. Optionally, each training sample includes a sample pronunciation feature, an emotion category label and a sample text. Alternatively, each training sample includes a sample pronunciation feature, an emotion category label, a sample text and a speech label.

[0094] Optionally, the machine learning model used to train the speech synthesis model includes, but is not limited to, a transformer model, a convolutional neural network (CNN), a recurrent neural network (RNN) or a deep neural network (DNN). The embodiments of the present application do not limit the training method and training samples of the speech synthesis model.

[0095] In the case where each training sample includes a sample pronunciation feature, an emotion category label and a sample text, during training, the training sample is input into the machine learning model to obtain an output speech of the machine learning model. The output speech is subjected to pronunciation feature extraction and emotion recognition to obtain the pronunciation feature and the emotion category corresponding to the output speech. It is determined whether the emotion category corresponding to the output speech is consistent with the emotion category label, and the cross entropy between the pronunciation feature corresponding to the output speech and the sample pronunciation feature is calculated. The machine learning model is iteratively trained until a preset accuracy requirement is met.

[0096] In the case where each training sample includes a sample pronunciation feature, an emotion category label, a sample text and a speech label, during training, the training sample is input into the machine learning model to obtain an output speech of the machine learning model. The mean square error between the output speech and the speech label is calculated. The machine learning model is iteratively trained until a preset accuracy requirement is met.

[0097] In step 104, the computing device outputs a simultaneous interpretation result corresponding to the first speech, and the simultaneous interpretation result includes the second speech.

[0098] Optionally, in the case where the method 100 is applied to a conference system, the first speech is from a first conference terminal in the conference system. In this case, the implementation of the computing device outputting the simultaneous interpretation result corresponding to the first speech includes: the computing device sending the simultaneous interpretation result to a second conference terminal in the conference system.

[0099] Optionally, before performing the above step 102, the computing device receives a language selection instruction sent by the second conference terminal, the language selection instruction being used to indicate selection of the second language. The computing device translates the first text in the first language into the second text in the second language based on the language selection instruction.

[0100] In the embodiments of the present application, the computing device acquires the pronunciation feature and the emotion category corresponding to the original speech, and in the process of translating the original speech into the target speech in the specified language, the pronunciation feature and the emotion category corresponding to the original speech are imitated by the machine, so that the translation result is output in the pronunciation feature and the emotion category corresponding to the original speech, that is, the pronunciation feature of the target speech is consistent with that of the original speech and the reflected emotion state is the same, so that the translation result is more authentic and the translation effect is improved. In the cross-language communication scenario, the communication effect of the same language is realized, thereby improving the user experience.

[0101] Optionally, the above method 100 is performed by a simultaneous interpretation engine in the computing device. For example, Figure 2 is a functional module schematic diagram of a simultaneous interpretation engine provided by the embodiments of the present application. As shown in the figure, Figure 2 the simultaneous interpretation engine includes a speech recognition module, a machine translation module, a speech synthesis module, an emotion recognition module, a voiceprint recognition module, and a feature matching module. The speech recognition module is used to convert the original speech into the original text. The machine translation module is used to translate the original text into the target text in the target language, and input the target text into the speech synthesis module. The emotion recognition module is used to perform emotion recognition on the original speech to obtain the emotion category corresponding to the original speech, and input the emotion category corresponding to the original speech into the speech synthesis module. The voiceprint recognition module is used to perform voiceprint recognition on the original speech to obtain the voiceprint feature corresponding to the original speech. The feature matching module is used to match the voiceprint feature corresponding to the original speech with the voiceprint features in the pronunciation feature set library. If the pronunciation feature set library includes the voiceprint feature corresponding to the original speech, the pronunciation feature corresponding to the voiceprint feature corresponding to the original speech in the pronunciation feature set library is acquired and input into the speech synthesis module. Optionally, the simultaneous interpretation engine further includes a pronunciation feature extraction module. If the pronunciation feature set library does not include the voiceprint feature corresponding to the original speech, the pronunciation feature extraction module is used to extract the pronunciation feature of the original speech to obtain the pronunciation feature corresponding to the original speech, and input the pronunciation feature corresponding to the original speech into the speech synthesis module, and store the voiceprint feature corresponding to the original speech and the pronunciation feature corresponding to the original speech in the pronunciation feature set library. The speech synthesis module is used to generate and output the target speech (including pronunciation feature and emotion) according to the input pronunciation feature corresponding to the original speech, the emotion category corresponding to the original speech, and the target text. It is worth noting that the pronunciation feature set library can be deployed in the simultaneous interpretation engine, or it can be deployed independently of the simultaneous interpretation engine.

[0102] For another example, Figure 3 is a flowchart of another simultaneous interpretation method provided by an embodiment of the present application. The method 300 outputs audio and video. As shown in the figure, the method 300 includes but is not limited to the following steps 301 to 305. Figure 3

[0103] In step 301, the computing device obtains the character feature corresponding to the first voice, the emotion category corresponding to the first voice, and the first text corresponding to the first voice. The character feature includes pronunciation feature and character image, and the character image includes the face of the target character.

[0104] The first text is used to describe the voice content of the first voice. In this step 301, the implementation manner of the computing device obtaining the character feature corresponding to the first voice, the emotion category corresponding to the first voice, and the first text corresponding to the first voice can refer to the above-mentioned step 101 respectively, and the embodiments of the present application will not be repeated here.

[0105] Compared with the above-mentioned method 100, the difference of this method 300 is that the character feature also includes the character image. Therefore, in this method 300, the target feature set library includes the corresponding relationship between multiple sets of voiceprint features, pronunciation features, and character images. For example, the target feature set library stores the corresponding relationship between multiple sets of voiceprint features, pronunciation features, and character images in the form of key-value pairs, taking voiceprint features as keys, and pronunciation features and character images as values. The target feature set library can be represented as follows: {voiceprint feature 1: pronunciation feature 1, character image 1; voiceprint feature 2: pronunciation feature 2, character image 2; voiceprint feature 3: pronunciation feature 3, character image 3; ……}.

[0106] ​Optionally, the manner of obtaining the correspondence between the voiceprint feature and the pronunciation feature in the target feature set library can refer to step 1012 described above. The character image can be pre-acquired from a recording video, a photographed image or a digital human provided by the user during user registration. For example, the computing device acquires user registration information including a recording voice and a photographed image including the target character. The computing device performs voiceprint recognition on the recording voice to obtain the voiceprint feature corresponding to the recording voice, performs pronunciation feature extraction on the recording voice to obtain the pronunciation feature corresponding to the recording voice, and performs character image acquisition on the photographed image to obtain the character image of the target character. The computing device correspondingly stores the voiceprint feature corresponding to the recording voice, the pronunciation feature corresponding to the recording voice and the character image of the target character in the target feature set library. Alternatively, the character image can be acquired by image acquisition on the user during user speech. For example, when the voiceprint feature of a speaker is not included in the target feature set library, the pronunciation feature of the speaker is extracted from the speech content of the speaker to obtain the pronunciation feature of the speaker, the character image of the speaker is acquired to obtain the character image of the speaker, and the voiceprint feature, the pronunciation feature and the character image of the speaker are further correspondingly stored in the target feature set library. Subsequently, when the speaker speaks again, since the voiceprint feature of the speaker has been included in the target feature set library, the computing device does not need to repeatedly extract the pronunciation feature of the speaker and acquire the character image of the speaker, thereby improving the simultaneous interpretation efficiency. Alternatively, in the case that no information for extracting the character image is provided during user registration and the character image of the user cannot be acquired during user speech, the computing device can use a default character image.

[0107] Step 302: The computing device translates the first text into a second text.

[0108] The first text is in a first language, and the second text is in a second language different from the first language. The implementation of step 302 can refer to step 102 described above, and details are not repeated here.

[0109] Step 303: The computing device generates a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice and the second text.

[0110] The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice, and the emotional state reflected by the second voice is the same as the emotional state reflected by the first voice. The implementation of step 303 can refer to step 203 described above, and details are not repeated here.

[0111] At step 304, the computing device performs speech driving on the character image using the second speech to obtain a character video in which the mouth movement of the target character matches the second speech.

[0112] Speech driving is a data-driven, artificial intelligence face animation generation algorithm. The input of the algorithm is usually a face video and a speech audio, and the main goal of the algorithm is to edit the mouth of the face in the input video to match the input speech audio.

[0113] Optionally, the character image is provided in the form of an original video. The computing device performs speech driving on the original video using the second speech, that is, changes the mouth movement of the target character in the original video according to the second speech, so that the mouth movement of the target character in the final obtained character video matches the second speech. The implementation process of this scheme is as follows. First, the face in the original video is detected and cropped, and the lower half of the face (including the mouth region) is further masked, and a face that has not been masked is randomly selected as a reference frame (the part of the face that has not been masked is used to provide appearance and pose information to ensure better generation effect). The masked face and the face as the reference frame are used as face input information. Second, the second speech is cut, sampled and feature extracted, and the matched mouth shape information is obtained according to the extracted audio features. Finally, the face input information and the mouth shape information are input into the pre-trained generation network to perform mouth synthesis. The picture effect of the final obtained character video is equivalent to changing the mouth shape of the character on the basis of the original video to make it consistent with the speech content.

[0114] Alternatively, the character image is provided in the form of a single picture. The computing device performs speech driving on the single picture using the second speech to generate a character video. The implementation process of this scheme is as follows. First, the second speech is cut, sampled and feature extracted, and the matched mouth shape information, head pose information and expression information are obtained according to the extracted audio features. Then, the single picture and the mouth shape information, head pose information and expression information are input into the pre-trained generation network for synthesis driving, and a character video is output. The picture effect of each frame of image in the final output character video is equivalent to changing the mouth shape, pose or expression of the character on the basis of the single picture.

[0115] It is worth noting that the above two speech driving schemes (for original video or single picture) are only used as examples, and the speech driving technology used by the computing device is not limited by the embodiments of the present application. Optionally, in the process of performing speech driving on the character image using the second speech, the expression and / or body movement of the character image can also be arranged and rendered according to the speech content and the emotional category, so that the expression and body movement of the target character in the character video are more matched with the second speech and the emotional expression of the speaker.

[0116] At step 305, the computing device outputs a simultaneous interpretation result corresponding to the first speech, the simultaneous interpretation result including the second speech and the video of the character.

[0117] In the embodiments of the present application, the computing device acquires the pronunciation feature and the emotion category corresponding to the original speech, and in the process of translating the original speech into the target speech in the specified language, the pronunciation feature and the emotion category corresponding to the original speech are imitated by the machine, so that the translation result is output in the pronunciation feature and the emotion category corresponding to the original speech, that is, the target speech is consistent with the pronunciation feature corresponding to the original speech and the emotion state reflected is the same, so that the translation result is more authentic, and the translation effect is improved. In the cross-language communication scene, the communication effect of the same language is realized, thereby improving the user experience. In addition, the computing device outputs the video of the character whose mouth movement matches the target speech, so that when the target speech and the video of the character are played, the target speech and the mouth shape of the character in the video of the character are consistent, thereby realizing the synchronous playing of the audio and the video, and improving the user experience. In the case where the character image is a real person image, such as the case where the character image is the image of the speaker providing the first speech, the present application can provide a high-fidelity audio and video of the speaker speaking in a non-native language (the second language) for the user to watch, thereby improving the user experience.

[0118] Optionally, the method 300 is executed by a simultaneous interpretation engine in the computing device. For example, Figure 4 is another functional module schematic diagram of the simultaneous interpretation engine provided by the embodiments of the present application. As shown in Figure 4As shown, the simultaneous interpretation engine includes a speech recognition module, a machine translation module, a speech synthesis module, an emotion recognition module, a voiceprint recognition module, a feature matching module, and a video generation module. The speech recognition module is configured to convert the original speech into original text. The machine translation module is configured to translate the original text into target text in a target language, and input the target text into the speech synthesis module. The emotion recognition module is configured to perform emotion recognition on the original speech to obtain an emotion category corresponding to the original speech, and input the emotion category corresponding to the original speech into the speech synthesis module. The voiceprint recognition module is configured to perform voiceprint recognition on the original speech to obtain a voiceprint feature corresponding to the original speech. The feature matching module is configured to match the voiceprint feature corresponding to the original speech with voiceprint features in a target feature set library. If the target feature set library includes the voiceprint feature corresponding to the original speech, the feature matching module is configured to obtain a pronunciation feature corresponding to the voiceprint feature corresponding to the original speech in the target feature set library, and input the pronunciation feature into the speech synthesis module, and / or obtain a character image corresponding to the voiceprint feature corresponding to the original speech in the target feature set library, and input the character image into the video generation module. Optionally, the simultaneous interpretation engine further includes a pronunciation feature extraction module and / or a character image extraction module. If the target feature set library does not include the pronunciation feature corresponding to the voiceprint feature corresponding to the original speech, the pronunciation feature extraction module is configured to perform pronunciation feature extraction on the original speech to obtain a pronunciation feature corresponding to the original speech, and input the pronunciation feature corresponding to the original speech into the speech synthesis module, and store the voiceprint feature corresponding to the original speech and the pronunciation feature corresponding to the original speech in the target feature set library correspondingly. If the target feature set library does not include the character image corresponding to the voiceprint feature corresponding to the original speech, the character image extraction module is configured to perform character image extraction on an input video or picture to obtain a character image corresponding to the original speech, and input the character image corresponding to the original speech into the video generation module, and store the voiceprint feature corresponding to the original speech and the character image corresponding to the original speech in the target feature set library correspondingly. The speech synthesis module is configured to generate and output target speech according to the pronunciation feature corresponding to the original speech, the emotion category corresponding to the original speech, and the target text. The video generation module is configured to perform voice driving on the character image input by using the target speech to obtain a character video, and output the character video. Finally, the speech synthesis module and the video generation module jointly output audio and video. It should be noted that the target feature set library can be deployed in the simultaneous interpretation engine, or can be independently deployed separately from the simultaneous interpretation engine.

[0119] Optionally, in the case where the target feature set library does not include the pronunciation feature or the character image corresponding to the voiceprint feature corresponding to the original speech, a default pronunciation feature or character image can also be used, which is not limited in the embodiments of the present application.

[0120] The application of the scheme of the embodiments of the present application will be illustrated below in combination with a video conference scenario.

[0121] The video conference scenario in the embodiments of the present application is a multi-person multi-point conference, that is, the video conference scenario includes multiple conference participants, and one conference participant joins the video conference through one conference terminal. The form of the conference terminal can be a dedicated physical device or a software program with conference functions, which can run on various computing devices such as mobile phones, tablets, computers and various user terminals. In this case, the computing device running the software program can also be considered as a conference terminal. The conference terminal joins the video conference through a conference service platform. Specifically, the conference terminal can obtain media data of the video conference from the conference service platform, and send locally collected media data to the conference service platform, so that the conference service platform forwards it to other conference terminals of the conference participants. The conference terminals can be connected through a wireless network, so that the conference participants are not limited by geographical location and can successfully join the video conference. In some cases, one conference participant can include only one conference user, such as a conference user who joins the video conference by running a conference software program on a personal mobile phone. In this case, the conference terminal used by the conference participant can be referred to as a user terminal. In some cases, one conference participant can also include multiple conference users, such as multiple conference users in a conference room who join the video conference through a conference terminal in the conference room. In this case, the conference terminal used by the conference participant can be referred to as a conference room terminal. In the video conference scenario provided in the embodiments of the present application, a user terminal and / or a conference room terminal can access the conference. In the case where a conference room terminal accesses the conference, the conference room terminal supports single language or multiple languages to access the conference.

[0122] For example, Figure 5 is a schematic diagram of a video conference scenario provided in the embodiments of the present application. As Figure 5 shown, the video conference scenario includes a conference service platform (also referred to as a conference media server) and four conference participants (conference participant A, conference participant B, conference participant C and conference participant D). Among them, conference participant A includes conference terminal A and user A, that is, user A accesses the conference through conference terminal A. Conference participant B includes conference terminal B and user B, that is, user B accesses the conference through conference terminal B. Conference participant C includes conference terminal C and user C, that is, user C accesses the conference through conference terminal C. Conference participant D includes conference terminal D and users D1, D2 and D3, that is, users D1, D2 and D3 jointly access the conference through conference terminal D.

[0123] Optionally, the conference service platform is a multipoint control unit (MCU). The MCU can provide registration service and authentication service (through the conference terminal used by the conference participant) and forwarding of conference data. For example, conference terminal A sends the video data of user A side to the MCU, and the MCU forwards the video data to conference terminal B, conference terminal C and conference terminal D, so that user B, user C and user D1-D3 can watch the audio and video image of user A through conference terminal B, conference terminal C and conference terminal D respectively.

[0124] Optionally, please continue to refer to Figure 5 The video conference scene further includes a simultaneous interpretation server. Optionally, the simultaneous interpretation server is a server, or a server cluster composed of multiple servers, or a cloud service platform. The simultaneous interpretation server stores a target feature set library. The target feature set library is the target feature set library in the method 100 or the method 300. Optionally, one or more simultaneous interpretation engines are deployed in the simultaneous interpretation server. Alternatively, the simultaneous interpretation engine can also be deployed in the conference terminal.

[0125] Optionally, after the user completes the registration in the conference service platform, the user logs in and accesses the conference as a registered user. In the new user registration stage, the user can provide character information, which includes recording voice, so as to extract the voiceprint features and pronunciation features of the user and store them in the target feature set library. Optionally, the character information can also include recording video or taking pictures, so as to provide the character image of the user and store the character image and the voiceprint features of the user in the target feature set library. For example, Figure 6 is a display interface schematic diagram provided by an embodiment of the present application. The display interface is a conference registration interface displayed on the conference terminal. As shown in Figure 6 The registration control M is displayed in the conference registration interface, and the registration control M includes an account input box, a password input box, a character information input option and a confirmation option. The account input box is used to input the registered account. The password input box is used to input the registered password set by the user. The character information input option is used for the user to provide character information. For example, the character information input option includes a voice input option and an image input option, the voice input option is used for the user to provide recorded voice, and the image input option is used for the user to provide recorded video or take pictures. Alternatively, the user can also access the conference as a visitor (unregistered user) temporarily.

[0126] The embodiments of the present application can be applied to various video conference scenarios. For example, scenario 1: a cross-language international conference with only user terminals accessing. Before the online cross-language international conference starts, as a Chinese, Xiaoa first needs to register as a new user. During the registration, Xiaoa needs to record a voice. The conference system extracts the voiceprint features and pronunciation features of Xiaoa and saves them in the target feature set library. Similarly, other users can also choose to record a voice after accessing the conference. The conference system saves the voiceprint features and pronunciation features of the user who records the voice. When the conference starts, Xiaoa accesses the conference according to the conference link. After accessing the conference, Xiaoa can select the Chinese language that Xiaoa wants to receive through the conference interface. When a British user (user B) speaks in English, the audio heard by Xiaoa is the cloned audio of user B after translation. During the conference, a German user (user C) temporarily accesses the conference as a visitor. When user C speaks, the conference system extracts the pronunciation features of user C. The audio heard by Xiaoa is the cloned audio of user C after translation. Scenario 2: a cross-language international conference with both user terminals and conference room terminals of the same language accessing. Scenario 2 is similar to scenario 1. Because the conference room participants use the same language, only one language needs to be selected. When the conference room participants speak, the system extracts the pronunciation features of the speakers and saves them. Scenario 3: a cross-language international conference with both user terminals and conference room terminals of different languages accessing. In scenario 3, because the conference room has participants who use different languages, these participants can choose to wear earphones to participate in the conference. After selecting a language, each participant listens to the translated audio of other participants through the earphones.

[0127] Optionally, the method 100 or the method 300 can be performed by different roles in the conference system or can be performed by multiple roles in the conference system in cooperation.

[0128] In the first case, the method 100 or the method 300 is performed by a simultaneous interpretation server in the conference system, that is, a simultaneous interpretation engine is deployed in the simultaneous interpretation server. For specific implementation, reference can be made to the first optional embodiment and the second optional embodiment.

[0129] In the first optional embodiment of the present application, the simultaneous interpretation server is connected to the conference service platform. For example, Figure 7 is a structural schematic diagram of a conference system provided by an embodiment of the present application. As Figure 7 indicated, the conference system includes a simultaneous interpretation server, a conference service platform, and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The conference service platform is connected to the simultaneous interpretation server. The plurality of conference terminals include a first conference terminal and a second conference terminal. In the embodiments of the present application, the first conference terminal is taken as a speaker terminal, and the second conference terminal is taken as a listener terminal.

[0130] The conference service platform is configured to receive the first voice sent by the first conference terminal and send the first voice to the simultaneous interpretation server. The simultaneous interpretation server is configured to obtain a character feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text. The character feature includes a pronunciation feature, the first text is used to describe the voice content of the first voice, the first text is in a first language, the second text is in a second language, and the second language is different from the first language. The simultaneous interpretation server is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform, the simultaneous interpretation result including the second voice, the voice content of the second voice being description content of the second text, the pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and the emotion state reflected by the second voice being the same as the emotion state reflected by the first voice. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

[0131] Optionally, the first conference terminal is configured to display a conference registration interface, and the conference registration interface displays a character information input option, the character information input option being used for a user to provide character information, the character information including a recorded voice. The first conference terminal is further configured to generate user registration information in response to receiving the recorded voice through the conference registration interface, the user registration information including the recorded voice, and the recorded voice being used to determine a corresponding relationship between a group of voiceprint features and pronunciation features. Further, the first conference terminal is further configured to send the user registration information to the conference service platform. The conference service platform is configured to send the user registration information to the simultaneous interpretation server. The simultaneous interpretation server is configured to extract and store the voiceprint features and the pronunciation features of the user according to the user registration information.

[0132] Optionally, the second conference terminal is configured to receive a language selection instruction through a display interface, the language selection instruction being used to indicate selection of the second language. The second conference terminal is configured to send the language selection instruction to the conference service platform. The conference service platform is configured to send the language selection instruction to the simultaneous interpretation server. The simultaneous interpretation server is configured to translate the first text into the second text according to the selected second language.

[0133] For example, refer to Figure 7 The conference system shown, the simultaneous interpretation server is deployed with a simultaneous interpretation engine, and the function modules in the simultaneous interpretation engine can refer to Figure 2 or Figure 4The implementation process of the scheme of the embodiment of the application is as follows: 1, the speaker uses the tool provided or specified by the simultaneous interpretation server to input own voice, video or photo, the simultaneous interpretation server performs machine learning and inference preprocessing, and completes digital modeling of the voice and image of the speaker. 2, after the video conference is held, the conference service platform automatically generates corresponding language channels and an original sound channel according to the preset language, and the speaker's voice is transmitted to the original sound channel. 3, the conference service platform transmits the voice in the original sound channel to the simultaneous interpretation server for processing. 4, the simultaneous interpretation server uses voiceprint recognition to obtain corresponding audio and video digital modeling data (i.e. character features), obtains translated target text through voice recognition and machine translation. Combined with the audio modeling data, the corresponding cloned voice is generated, and the cloned voice is used to drive the inference and rendering generation of the video, and finally the audio and video synchronization processing is completed. 5, the simultaneous interpretation server returns the real-time generated audio and video to the corresponding language channel of the conference service platform. 6, the conference service platform pushes the audio and video generated by the simultaneous interpretation server according to the language channel selected by the listener terminal.

[0134] In Figure 7 The conference system shown in the figure, the user through the requirement of the simultaneous interpretation server to input own voice, video or photo for large model inference learning to complete modeling. The voice of the conference speaker is forwarded to the simultaneous interpretation server by the conference service platform, the simultaneous interpretation server obtains the corresponding model data according to the voiceprint, synchronously translates and generates cloned voice, and generates corresponding audio and video according to the cloned voice, and returns the audio and video to the conference service platform. The conference service platform forwards the corresponding audio and video generated by the simultaneous interpretation server according to the language channel selected by other conference participants.

[0135] In the second optional embodiment of the application, the virtual interpreter terminal is connected to the conference service platform and the simultaneous interpretation server respectively, so that the conference service platform can focus on the conference business function, and avoid that a large number of AI control interactions cause the system to be too complex. For example, Figure 8 is another structure diagram of a conference system provided by the embodiment of the application. As Figure 8 shown, the conference system includes a simultaneous interpretation server, a conference service platform, a virtual interpreter terminal and a plurality of conference terminals, the plurality of conference terminals communicate through the conference service platform. The virtual interpreter terminal is connected to the conference service platform and the simultaneous interpretation server respectively. The plurality of conference terminals include a first conference terminal and a second conference terminal. The embodiment of the application takes the first conference terminal as the speaker terminal and the second conference terminal as the listener terminal as an example.

[0136] The conference service platform is configured to receive the first voice sent by the first conference terminal and send the first voice to the virtual interpreter terminal. The virtual interpreter terminal is configured to send the first voice to the simultaneous interpretation server. The simultaneous interpretation server is configured to obtain a person feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text. The person feature includes a pronunciation feature. The first text is used to describe the voice content of the first voice. The first text is in a first language. The second text is in a second language. The second language is different from the first language. The simultaneous interpretation server is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the virtual interpreter terminal. The simultaneous interpretation result includes the second voice. The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice. The emotion state reflected by the second voice is the same as the emotion state reflected by the first voice. The virtual interpreter terminal is configured to send the simultaneous interpretation result to the conference service platform. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

[0137] Optionally, the first conference terminal is configured to display a conference registration interface. The conference registration interface displays a person information input option. The person information input option is used for a user to provide person information. The person information includes a recorded voice. The first conference terminal is further configured to generate user registration information in response to receiving the recorded voice through the conference registration interface. The user registration information includes the recorded voice. The recorded voice is used to determine the correspondence between a group of voiceprint features and pronunciation features. Further, the first conference terminal is further configured to send the user registration information to the conference service platform. The conference service platform is configured to send the user registration information to the virtual interpreter terminal. The virtual interpreter terminal is configured to send the user registration information to the simultaneous interpretation server. The simultaneous interpretation server is configured to extract and store the voiceprint features and pronunciation features of the user according to the user registration information.

[0138] Optionally, the second conference terminal is configured to receive a language selection instruction through a display interface. The language selection instruction is used to indicate the selection of the second language. The second conference terminal is configured to send the language selection instruction to the conference service platform. The conference service platform is configured to send the language selection instruction to the virtual interpreter terminal. The virtual interpreter terminal is configured to send the language selection instruction to the simultaneous interpretation server. The simultaneous interpretation server is configured to translate the first text into the second text according to the selected second language.

[0139] For example, refer to Figure 8 The conference system shown, the simultaneous interpretation server is deployed with a simultaneous interpretation engine. The function modules in the simultaneous interpretation engine can refer to Figure 2 or Figure 4The implementation process of this application embodiment is as follows: 1. The speaker uses tools provided or specified by the simultaneous interpretation server to record their voice, video, or photos. The simultaneous interpretation server performs machine learning and inference preprocessing to complete the digital modeling of the speaker's voice and image. 2. After the video conference is held, the conference service platform automatically generates a corresponding language channel and an original audio channel according to the preset language. The speaker's voice is transmitted to the original audio channel. 3. The conference service platform transmits the audio from the original audio channel to the virtual interpreter terminal. 4. The virtual interpreter terminal transmits the audio from the original audio channel to the simultaneous interpretation server for processing. 5. The simultaneous interpretation server uses voiceprint recognition to obtain the corresponding audio and video digital modeling data (i.e., personal features), and obtains the translated target text through speech recognition and machine translation. Combined with the audio modeling data, it generates the corresponding cloned speech, and then uses the cloned speech to drive the inference and rendering of the video, finally completing the audio and video synchronization processing. 6. The simultaneous interpretation server transmits the real-time generated audio and video back to the virtual interpreter terminal. One virtual interpreter terminal can receive audio and video in multiple languages ​​simultaneously. 7. The virtual interpreter terminal transmits the audio and video back to the corresponding language channel on the conference service platform. 8. The conference service platform pushes the audio and video generated by the simultaneous interpretation server according to the language channel selected by the listener's terminal.

[0140] exist Figure 8 In the illustrated conference system, users input their voice, video, or photos at the request of the simultaneous interpretation server for large-scale model inference and learning to complete the modeling process. The speaker's voice is forwarded by the conference service platform to a virtual interpreter terminal, which connects to the simultaneous interpretation server without the need for a real interpreter. The virtual interpreter terminal transmits the speaker's voice to the simultaneous interpretation server, which retrieves the corresponding model data based on the voiceprint, synchronously translates and clones the voice, and generates corresponding audio and video based on the cloned voice, sending it back to the virtual interpreter terminal. The virtual interpreter terminal then transmits the audio and video to the conference service platform according to the corresponding language. The conference service platform forwards the corresponding audio and video generated by the simultaneous interpretation server according to the language channels selected by other participants. A single virtual interpreter terminal can support multiple audio and video streams, or multiple virtual interpreters can be configured to share the load.

[0141] Optionally, please continue to see Figure 7 or Figure 8 In such Figure 7 or Figure 8 In the conference system shown, the simultaneous interpretation server can deploy multiple simultaneous interpretation engines. Each simultaneous interpretation engine corresponds to a conference terminal in the conference system, or each simultaneous interpretation engine corresponds to a conference held through the conference system, or each simultaneous interpretation engine corresponds to a language, which is either the language to be translated or the target language.

[0142] In the embodiments of the present application, the simultaneous interpretation engines in the simultaneous interpretation server are deployed according to conference terminals, conference sites or language categories, one simultaneous interpretation engine provides simultaneous interpretation services for one conference terminal, one conference or one language category, and the problems of insecurity, complex routing, vulnerability to attack, string interpretation, poor disaster recovery and mutual interference of multi-path requests when one simultaneous interpretation engine corresponds to multiple conference terminals, multiple conferences or multiple languages are solved, thereby improving the security, flexibility and practicability of the simultaneous interpretation engine.

[0143] In the second case, the method 100 or the method 300 is executed by a conference terminal in a conference system, that is, a simultaneous interpretation engine is deployed in the conference terminal, and the specific implementation can refer to the third optional embodiment and the fourth optional embodiment below.

[0144] In the third optional embodiment of the present application, the method 100 or the method 300 is executed by a listener terminal. For example, Figure 9 is another structure schematic diagram of a conference system provided by the embodiments of the present application. As shown in Figure 9 The conference system includes a conference service platform and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The plurality of conference terminals include a first conference terminal and a second conference terminal. The embodiments of the present application take the first conference terminal as a speaker terminal and the second conference terminal as a listener terminal as an example.

[0145] The conference service platform is configured to receive the first voice sent by the first conference terminal and send the first voice to the second conference terminal. The second conference terminal is configured to obtain a character feature corresponding to the first voice, an emotion category corresponding to the first voice and a first text corresponding to the first voice, and translate the first text into a second text, the character feature includes a pronunciation feature, the first text is used to describe the voice content of the first voice, the first text is in a first language, and the second text is in a second language, the second language is different from the first language. The second conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice and the second text, and play a simultaneous interpretation result corresponding to the first voice, the simultaneous interpretation result includes the second voice, the voice content of the second voice is the description content of the second text, the pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice, and the emotion state reflected by the second voice is the same as the emotion state reflected by the first voice.

[0146] Optionally, please continue to refer to Figure 9The conference system further includes a simultaneous interpretation server, the plurality of conference terminals are connected with the simultaneous interpretation server respectively, the simultaneous interpretation server stores a target feature set library, and the target feature set library includes a plurality of sets of corresponding relationships between voiceprint features and person features. The second conference terminal is configured to perform voiceprint recognition on the first voice to obtain target voiceprint features corresponding to the first voice, and send the target voiceprint features to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint features with voiceprint features in the target feature set library, and if the target voiceprint features are included in the target feature set library, send person features corresponding to the target voiceprint features in the target feature set library to the second conference terminal. The second conference terminal is configured to take the person features corresponding to the target voiceprint features sent by the simultaneous interpretation server as the person features corresponding to the first voice.

[0147] It should be noted that the second conference terminal can be a user terminal used by a user to access the conference, or can be a conference room terminal (all participants in the same conference room use the same language), or can be a headset or user terminal connected with the conference room terminal for listening to the translated language (participants in the same conference room use different languages).

[0148] Optionally, the first conference terminal is configured to display a conference registration interface, and the conference registration interface displays a person information input option, the person information input option being used by a user to provide person information, the person information including a recorded voice. The first conference terminal is further configured to generate user registration information in response to receiving the recorded voice through the conference registration interface, the user registration information including the recorded voice, and the recorded voice being used to determine a corresponding relationship between a set of voiceprint features and pronunciation features. Further, the first conference terminal is further configured to send the user registration information to the conference service platform. The conference service platform is configured to send the user registration information to the simultaneous interpretation server. The simultaneous interpretation server is configured to extract and store voiceprint features and pronunciation features of the user according to the user registration information.

[0149] Optionally, the second conference terminal is configured to receive a language selection instruction through a display interface, the language selection instruction being used to indicate selection of a second language. The second conference terminal is configured to translate the first text into a second text according to the selected second language.

[0150] For example, refer to Figure 9 The conference system shown in the figure, the listener terminal is deployed with a simultaneous interpretation engine. Optionally, the simultaneous interpretation server is deployed with a simultaneous interpretation engine. The function modules in the simultaneous interpretation engine can refer to Figure 2 or Figure 4The implementation process of the scheme of the embodiments of the present application is as follows: 1, the speaker uses the tool provided or specified by the simultaneous interpretation server to input his own voice, video or photo, and the simultaneous interpretation server performs machine learning and inference preprocessing to complete the digital modeling of the speaker's voice and image. 2, after the video conference is held, the speaker's voice is transmitted to the original sound channel. 3, the conference service platform transmits the voice in the original sound channel to the listener terminal. 4, the listener terminal uses voiceprint recognition to obtain corresponding audio and video digital modeling data (i.e. character features). 5, the listener terminal obtains the translated target text according to the selected language through voice recognition and machine translation. Combine the audio modeling data to generate the corresponding cloned voice, and then generate the corresponding audio and video through the cloned voice driving video inference and rendering, and finally complete the audio and video synchronization processing and output the corresponding language audio and video.

[0151] In Figure 9 The conference system shown in the figure, the user through the requirement of the simultaneous interpretation server to input his own voice, video or photo for large model inference learning to complete modeling. The voice of the conference speaker is forwarded to the listener terminal by the conference service platform, the corresponding model data is obtained by the listener terminal according to the voiceprint, and the cloned voice is generated according to the selected language and the cloned voice, and the corresponding audio and video are generated according to the cloned voice and output.

[0152] It is worth noting that the function of the simultaneous interpretation engine in the listener terminal is similar to that of the simultaneous interpretation engine in the simultaneous interpretation server, but it is usually limited by the processing performance of the conference terminal, and the listener terminal can only complete the low-precision simultaneous interpretation function. If high-precision simultaneous interpretation function is required, the listener terminal can interact with the simultaneous interpretation server through the simultaneous interpretation engine deployed by itself, and the high-precision simultaneous interpretation function can be completed by the simultaneous interpretation engine in the simultaneous interpretation server, and the generated audio and video can be transmitted back to the listener terminal.

[0153] In the fourth optional embodiment of the present application, the above-mentioned method 100 or method 300 is executed by the speaker terminal. For example, Figure 10 is another structure diagram of a conference system provided by the embodiments of the present application. As Figure 10 shown, the conference system includes a conference service platform and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The plurality of conference terminals include a first conference terminal and a second conference terminal. The embodiments of the present application take the first conference terminal as the speaker terminal and the second conference terminal as the listener terminal as an example.

[0154] The first conference terminal is configured to collect the first voice, obtain a character feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text. The character feature includes a pronunciation feature. The first text is used to describe the voice content of the first voice. The first text is in a first language, and the second text is in a second language. The second language is different from the first language. The first conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform. The simultaneous interpretation result includes the second voice. The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice. The emotion state reflected by the second voice is the same as the emotion state reflected by the first voice. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

[0155] Optionally, please continue to refer to Figure 10 The conference system further includes a simultaneous interpretation server. The first conference terminal is connected with the simultaneous interpretation server. The simultaneous interpretation server stores a target feature set library. The target feature set library includes a plurality of corresponding relationships between voiceprint features and character features. The first conference terminal is configured to perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice, and send the target voiceprint feature to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint feature with the voiceprint features in the target feature set library. If the target feature set library includes the target voiceprint feature, the simultaneous interpretation server sends a character feature corresponding to the target voiceprint feature in the target feature set library to the first conference terminal. The first conference terminal is configured to use the character feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the character feature corresponding to the first voice.

[0156] Optionally, the first conference terminal is configured to display a conference registration interface. The conference registration interface displays a character information input option. The character information input option is used for a user to provide character information. The character information includes a recorded voice. The first conference terminal is further configured to generate user registration information in response to receiving the recorded voice through the conference registration interface. The user registration information includes the recorded voice. The recorded voice is used to determine a corresponding relationship between a group of voiceprint features and pronunciation features. Further, the first conference terminal is further configured to send the user registration information to the conference service platform. The conference service platform is configured to send the user registration information to the simultaneous interpretation server. The simultaneous interpretation server is configured to extract and store the voiceprint features and pronunciation features of the user according to the user registration information.

[0157] Optionally, the second conference terminal is configured to receive a language selection instruction through the display interface, the language selection instruction being used to indicate selection of a second language. The second conference terminal is configured to send the language selection instruction to the conference service platform. The conference service platform is configured to send the language selection instruction to the first conference terminal. The first conference terminal is configured to translate the first text into a second text according to the selected second language.

[0158] For example, refer to Figure 10 The illustrated conference system, the speaker terminal is deployed with a simultaneous interpretation engine. Optionally, the simultaneous interpretation server is deployed with a simultaneous interpretation engine. The function modules in the simultaneous interpretation engine can refer to Figure 2 Or Figure 4 The implementation process of the scheme of the embodiments of the present application is as follows: 1, the speaker uses the tool provided or specified by the simultaneous interpretation server to input his own voice, video or photo, and the simultaneous interpretation server performs machine learning and inference preprocessing to complete the digital modeling of the speaker's voice and image. 2, after the video conference is held, the conference service platform automatically generates corresponding language channels and an original sound channel according to the preset language, and the speaker's voice is transmitted to the original sound channel. 3, the speaker terminal uses voiceprint recognition to obtain corresponding audio and video digital modeling data (i.e. character features). 4, the speaker terminal obtains the translated target text through voice recognition and machine translation. Combine the audio modeling data to generate corresponding cloned voice, and then generate the inference and rendering of the video driven by the cloned voice, and finally complete the audio and video synchronization processing. The speaker terminal transmits the real-time generated audio and video to the corresponding language channel of the conference service platform. 5, the conference service platform pushes the audio and video generated by the simultaneous interpretation server according to the language channel selected by the listener terminal.

[0159] In Figure 10 In the illustrated conference system, the user inputs his own voice, video or photo through the requirement of the simultaneous interpretation server for large model inference learning to complete modeling. After the speaker terminal collects the voice of the conference speaker, the speaker terminal synchronously translates and generates cloned voice according to the corresponding model data obtained by voiceprint, and generates corresponding audio and video according to the cloned voice and sends it to the conference service platform. The conference service platform forwards the corresponding audio and video generated by the simultaneous interpretation server according to the language channel selected by other conference participants.

[0160] It is worth noting that the simultaneous interpretation engine in the speaker terminal is similar to the simultaneous interpretation engine in the simultaneous interpretation server, but is usually limited by the processing performance of the conference terminal, and the speaker terminal can only complete low-precision simultaneous interpretation function. If high-precision simultaneous interpretation function is required, the speaker terminal can interact with the simultaneous interpretation server through the simultaneous interpretation engine deployed by itself, and the simultaneous interpretation engine in the simultaneous interpretation server can complete high-precision simultaneous interpretation function, and the generated audio and video can be transmitted back to the speaker terminal.

[0161] In the third case, the method 100 or the method 300 is executed by a conference terminal and a conference service platform in a conference system, and the conference terminal is deployed with a simultaneous interpretation engine. The specific implementation can refer to the fifth optional embodiment and the sixth optional embodiment below.

[0162] In the fifth optional embodiment of the present application, the method 100 or the method 300 is executed by a listener terminal and a conference service platform. For example, Figure 11 is another structure diagram of a conference system provided by the embodiments of the present application. As shown in Figure 11 , the conference system includes a conference service platform and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The plurality of conference terminals include a first conference terminal and a second conference terminal. The plurality of conference terminals include a first conference terminal and a second conference terminal. The embodiments of the present application take the first conference terminal as the speaker terminal and the second conference terminal as the listener terminal as an example.

[0163] The conference service platform is configured to receive the first voice sent by the first conference terminal, obtain the character features corresponding to the first voice, and send the first voice and the character features corresponding to the first voice to the second conference terminal, wherein the character features include pronunciation features. The second conference terminal is configured to obtain the emotion category corresponding to the first voice and the first text corresponding to the first voice, and translate the first text into a second text, wherein the first text is used to describe the voice content of the first voice, the first text is in a first language, and the second text is in a second language, and the second language is different from the first language. The second conference terminal is further configured to generate a second voice according to the pronunciation features corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and play a simultaneous interpretation result of the first voice, wherein the simultaneous interpretation result includes the second voice, the voice content of the second voice is the description content of the second text, the pronunciation features corresponding to the second voice are consistent with the pronunciation features corresponding to the first voice, and the emotion state reflected by the second voice is the same as the emotion state reflected by the first voice.

[0164] Optionally, please continue to refer to Figure 11The conference system further includes a simultaneous interpretation server, the conference service platform is connected with the simultaneous interpretation server, the simultaneous interpretation server stores a target feature set library, and the target feature set library includes a correspondence relationship between a plurality of groups of voiceprint features and person features. The conference service platform is configured to perform voiceprint recognition on the first voice to obtain target voiceprint features corresponding to the first voice, and send the target voiceprint features to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint features with voiceprint features in the target feature set library, and if the target voiceprint features are included in the target feature set library, send person features corresponding to the target voiceprint features in the target feature set library to the conference service platform. The conference service platform is configured to take the person features corresponding to the target voiceprint features sent by the simultaneous interpretation server as the person features corresponding to the first voice. Alternatively, the target feature set library is stored in the conference service platform, and thus the simultaneous interpretation server does not need to be separately deployed. The conference service platform is configured to match the target voiceprint features with voiceprint features in the target feature set library, and if the target voiceprint features are included in the target feature set library, take person features corresponding to the target voiceprint features in the target feature set library as the person features corresponding to the first voice.

[0165] It should be noted that the second conference terminal can be a user terminal used by a user to access the conference, or can be a conference room terminal (all participants in the same conference room use the same language), or can be a headset or user terminal connected to the conference room terminal for listening to the translated language (participants in the same conference room use different languages).

[0166] Optionally, the first conference terminal is configured to display a conference registration interface, and the conference registration interface displays a person information input option, the person information input option being used by a user to provide person information, and the person information including a recorded voice. The first conference terminal is further configured to, in response to receiving the recorded voice through the conference registration interface, generate user registration information, the user registration information including the recorded voice, and the recorded voice being used to determine a correspondence relationship between a group of voiceprint features and pronunciation features. Further, the first conference terminal is further configured to send the user registration information to the conference service platform. The conference service platform is configured to send the user registration information to the simultaneous interpretation server. The simultaneous interpretation server is configured to extract and store voiceprint features and pronunciation features of the user in correspondence according to the user registration information.

[0167] Optionally, the second conference terminal is configured to receive a language selection instruction through a display interface, the language selection instruction being used to indicate selection of a second language. The second conference terminal is configured to translate the first text into a second text according to the selected second language.

[0168] For example, refer to Figure 11 The conference system shown in the figure, the listener terminal is deployed with a simultaneous interpretation engine. Optionally, the simultaneous interpretation server is deployed with a simultaneous interpretation engine. The function modules in the simultaneous interpretation engine can refer toFigure 2 Or Figure 4 The implementation process of the scheme of the embodiments of the present application is as follows: 1. The speaker uses the tool provided or specified by the simultaneous interpretation server to input own voice, video or photo, and the simultaneous interpretation server performs machine learning and inference preprocessing to complete digital modeling of the voice and image of the speaker. 2. After the video conference is held, the voice of the speaker is transmitted to the original sound channel. 3. The conference service platform uses voiceprint recognition to obtain corresponding audio and video digital modeling data. 4. The conference service platform transmits the voice in the original sound channel and the audio and video digital modeling data (i.e. the characteristics of the person) to the listener terminal for processing. 5. The listener terminal obtains the translated target text through voice recognition and machine translation according to the selected language. The corresponding cloned voice is generated in combination with the audio modeling data, and the inference and rendering of the video are generated through the cloned voice, and finally the audio and video synchronization processing is completed, and the audio and video of the corresponding language are output.

[0169] In Figure 11 The conference system shown in the figure, the user through the requirement of the simultaneous interpretation server to input own voice, video or photo for large model inference learning to complete modeling. The voice of the conference speaker is forwarded to the listener terminal by the conference service platform, and the corresponding model data is obtained by the conference service platform according to the voiceprint and provided to the listener terminal, and the listener terminal generates cloned voice according to the selected language and drives the generation of corresponding audio and video, and outputs the audio and video.

[0170] It is worth noting that the function of the simultaneous interpretation engine in the listener terminal is similar to that of the simultaneous interpretation engine in the simultaneous interpretation server, but it is usually limited by the processing performance of the conference terminal, and the listener terminal can only complete the low-precision simultaneous interpretation function. If high-precision simultaneous interpretation function is needed, the listener terminal can interact with the simultaneous interpretation server through the simultaneous interpretation engine deployed by itself, and the high-precision simultaneous interpretation function is completed by the simultaneous interpretation engine in the simultaneous interpretation server, and the generated audio and video are transmitted back to the listener terminal.

[0171] In the sixth optional embodiment of the present application, the above-mentioned method 100 or method 300 is executed by the speaker terminal and the conference service platform. For example, Figure 12 is another structure diagram of a conference system provided by the embodiments of the present application. As Figure 12 shown, the conference system includes a conference service platform and a plurality of conference terminals. The plurality of conference terminals communicate through the conference service platform. The plurality of conference terminals include a first conference terminal and a second conference terminal. The embodiments of the present application take the first conference terminal as the speaker terminal and the second conference terminal as the listener terminal as an example.

[0172] The first conference terminal is configured to collect the first voice and send the first voice to the conference service platform. The conference service platform is configured to obtain a character feature corresponding to the first voice, and send the character feature corresponding to the first voice to the second conference terminal. The character feature includes a pronunciation feature. The first conference terminal is configured to obtain an emotion category corresponding to the first voice and a first text corresponding to the first voice, and translate the first text into a second text. The first text is used to describe the voice content of the first voice. The first text is in a first language, and the second text is in a second language. The second language is different from the first language. The first conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform. The simultaneous interpretation result includes the second voice. The voice content of the second voice is the description content of the second text. The pronunciation feature corresponding to the second voice is consistent with the pronunciation feature corresponding to the first voice. The emotion state reflected by the second voice is the same as the emotion state reflected by the first voice. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

[0173] Optionally, please continue to refer to Figure 12 The conference system further includes a simultaneous interpretation server. The conference service platform is connected with the simultaneous interpretation server. The simultaneous interpretation server stores a target feature set library. The target feature set library includes a plurality of corresponding relationships between voiceprint features and character features. The conference service platform is configured to perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice, and send the target voiceprint feature to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint feature with voiceprint features in the target feature set library. If the target voiceprint feature is included in the target feature set library, the character feature corresponding to the target voiceprint feature in the target feature set library is sent to the conference service platform. The conference service platform is configured to take the character feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the character feature corresponding to the first voice. Alternatively, the conference service platform stores the target feature set library. Therefore, the simultaneous interpretation server does not need to be separately deployed. The conference service platform is configured to match the target voiceprint feature with voiceprint features in the target feature set library. If the target voiceprint feature is included in the target feature set library, the character feature corresponding to the target voiceprint feature in the target feature set library is taken as the character feature corresponding to the first voice.

[0174] Optionally, the first conference terminal is configured to display a conference registration interface, and the conference registration interface displays a character information input option, the character information input option being configured to receive character information from a user, the character information comprising a recorded voice. The first conference terminal is further configured to generate user registration information in response to receiving the recorded voice via the conference registration interface, the user registration information comprising the recorded voice, and the recorded voice being configured to determine a corresponding relationship between a set of voiceprint features and pronunciation features. Further, the first conference terminal is further configured to send the user registration information to the conference service platform. The conference service platform is configured to send the user registration information to the simultaneous interpretation server. The simultaneous interpretation server is configured to extract and store the voiceprint features and the pronunciation features of the user according to the user registration information.

[0175] Optionally, the second conference terminal is configured to receive a language selection instruction via the display interface, the language selection instruction being configured to indicate a selected second language. The second conference terminal is configured to send the language selection instruction to the conference service platform. The conference service platform is configured to send the language selection instruction to the first conference terminal. The first conference terminal is configured to translate the first text into a second text according to the selected second language.

[0176] For example, refer to Figure 12 The conference system is shown, and the simultaneous interpretation engine is deployed in the speaker terminal. Optionally, the simultaneous interpretation engine is deployed in the simultaneous interpretation server. The function modules in the simultaneous interpretation engine can refer to Figure 2 or Figure 4 The implementation process of the scheme of the embodiments of the present application is as follows: 1. The speaker uses the tool provided or specified by the simultaneous interpretation server to input his own voice, video or photo, and the simultaneous interpretation server performs machine learning and inference preprocessing to complete the digital modeling of the speaker's voice and image. 2. After the video conference is held, the conference service platform automatically generates a corresponding language channel and an original sound channel according to the preset language, and the speaker's voice is transmitted to the original sound channel. 3. The conference service platform uses voiceprint recognition to obtain corresponding audio and video digital modeling data. 4. The conference service platform transmits the audio and video digital modeling data (i.e. character features) to the speaker terminal for processing. 5. The speaker terminal obtains the translated target text through voice recognition and machine translation. The corresponding cloned voice is generated in combination with the audio modeling data, and the cloned voice is used to drive the inference and rendering generation of the video, and finally the audio and video synchronization processing is completed. The speaker terminal transmits the real-time generated audio and video to the corresponding language channel of the conference service platform. 6. The conference service platform pushes the audio and video generated by the simultaneous interpretation server according to the language channel selected by the listener terminal.

[0177] In Figure 12 ​​In the illustrated conference system, a user enters his / her voice, video or photo for large model inference learning to complete modeling at the request of a simultaneous interpretation server. The conference service platform obtains corresponding model data according to the voiceprint and provides it to the speaker terminal. The speaker terminal generates translated and cloned voice according to the cloned voice, and generates corresponding audio and video according to the cloned voice and sends them to the conference service platform. The conference service platform forwards the corresponding audio and video generated by the simultaneous interpretation server to the language channel selected by other conference participants.

[0178] It is worth noting that the simultaneous interpretation engine in the speaker terminal has similar functions to the simultaneous interpretation engine in the simultaneous interpretation server, but is usually limited by the processing performance of the conference terminal. The speaker terminal can only complete low-precision simultaneous interpretation functions. If high-precision simultaneous interpretation functions are required, the speaker terminal can interact with the simultaneous interpretation server through its own deployed simultaneous interpretation engine, and the simultaneous interpretation engine in the simultaneous interpretation server can complete high-precision simultaneous interpretation functions and transmit the generated audio and video back to the speaker terminal.

[0179] The conference system provided by the embodiments of the present application can provide high-fidelity audio and video for conference speakers speaking in a non-native language for other conference participants to watch, achieving the effect of exchanging the same language in a cross-language exchange scenario. During the conference, the alternately speaking of different speakers can be recognized according to the voiceprint, and the pronunciation characteristics and emotional categories of the speakers can be imitated to generate and output the cloned voice of the speaker in the process of simultaneous interpretation, thereby improving the user experience. In addition, by deploying the simultaneous interpretation engine according to the number of conference terminals, the number of conference sessions or the number of language categories, one simultaneous interpretation engine provides simultaneous interpretation services for one conference terminal, one conference session or one language, solving the problems of insecurity, complex routing, vulnerability to attack, string translation, poor disaster recovery and mutual interference of multiple requests when one simultaneous interpretation engine corresponds to multiple conference terminals, multiple conference sessions or multiple languages, thereby improving the security, flexibility and practicality of the simultaneous interpretation engine.

[0180] The order of the steps of the simultaneous interpretation method provided by the embodiments of the present application can be adjusted appropriately, and the steps can also be increased or decreased accordingly. Any person skilled in the art can easily think of changes within the scope of the technology disclosed in the present application, which should be covered within the protection scope of the present application. For example, the deployment position of the simultaneous interpretation engine is not limited in the embodiments of the present application. In the conference system, if the simultaneous interpretation engine is deployed according to the number of conference terminals, the simultaneous interpretation engine can be deployed on the simultaneous interpretation server or the conference terminal. If the simultaneous interpretation engine is deployed according to the number of conference sessions or the number of language categories, the simultaneous interpretation engine can be deployed on the simultaneous interpretation server.

[0181] The virtual device of the embodiments of the present application is illustrated below.

[0182] For example, Figure 13 is a structural schematic diagram of a simultaneous interpretation device 1300 provided by an embodiment of the present application. As shown in Figure 13 , the simultaneous interpretation device 1300 includes but is not limited to an acquisition module 1301, a translation module 1302, a generation module 1303, and an output module 1304. Optionally, please continue to refer to Figure 13 , the simultaneous interpretation device 1300 further includes one or more of a storage module 1305, a voiceprint recognition module 1306, a voice driving module 1307, or a receiving module 1308.

[0183] The acquisition module 1301 is configured to acquire a character feature corresponding to a first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, the character feature including a pronunciation feature, and the first text being used to describe a voice content of the first voice.

[0184] The translation module 1302 is configured to translate the first text into a second text, the first text being in a first language, and the second text being in a second language, the second language being different from the first language.

[0185] The generation module 1303 is configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, a voice content of the second voice being a description content of the second text, a pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and an emotion state reflected by the second voice being the same as an emotion state reflected by the first voice.

[0186] The output module 1304 is configured to output a simultaneous interpretation result corresponding to the first voice, the simultaneous interpretation result including the second voice.

[0187] Optionally, the acquisition module 1301 is specifically configured to: perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice; match the target voiceprint feature with voiceprint features in a target feature set library, the target feature set library including a plurality of corresponding relationships between voiceprint features and character features; and if the target voiceprint feature is included in the target feature set library, take a character feature corresponding to the target voiceprint feature in the target feature set library as the character feature corresponding to the first voice.

[0188] Optionally, the acquisition module 1301 is further specifically configured to, if the target voiceprint feature is not included in the target feature set library, perform pronunciation feature extraction on the first voice to obtain the pronunciation feature corresponding to the first voice. The storage module 1305 is configured to store the target voiceprint feature and the pronunciation feature corresponding to the first voice in the target feature set library.

[0189] Optionally, the obtaining module 1301 is further configured to obtain user registration information, and the user registration information comprises the recorded voice. The voiceprint recognition module 1306 is configured to perform voiceprint recognition on the recorded voice to obtain voiceprint features corresponding to the recorded voice, and perform pronunciation feature extraction on the recorded voice to obtain pronunciation features corresponding to the recorded voice. The storage module 1305 is configured to store the voiceprint features corresponding to the recorded voice and the pronunciation features corresponding to the recorded voice in the target feature set library correspondingly.

[0190] Optionally, the character features further comprise a character image, and the character image comprises a face of the target character. The voice driving module 1307 is configured to perform voice driving on the character image by using the second voice to obtain a character video, and a mouth movement of the target character in the character video matches the second voice. The simultaneous interpretation result further comprises the character video.

[0191] Optionally, the obtaining module 1301 is specifically configured to perform emotion recognition on the first voice to obtain an emotion category corresponding to the first voice.

[0192] Optionally, the generating module 1303 is specifically configured to input the pronunciation features corresponding to the first voice, the emotion category corresponding to the first voice, and the second text into the voice synthesis model to obtain the second voice output by the voice synthesis model.

[0193] Optionally, the simultaneous interpretation apparatus 1300 is applied to a conference system, and the first voice is from a first conference terminal in the conference system. The output module 1304 is specifically configured to send the simultaneous interpretation result to a second conference terminal in the conference system.

[0194] Optionally, the receiving module 1308 is configured to receive a language selection instruction sent by the second conference terminal, and the language selection instruction is used to indicate selection of the second language.

[0195] The obtaining module 1301, the translation module 1302, the generating module 1303, the output module 1304, the storage module 1305, the voiceprint recognition module 1306, the voice driving module 1307, and the receiving module 1308 can be implemented by software or by hardware. For example, the implementation of the obtaining module 1301 is described below. Similarly, the implementation of the translation module 1302, the generating module 1303, the output module 1304, the storage module 1305, the voiceprint recognition module 1306, the voice driving module 1307, and the receiving module 1308 can refer to the implementation of the obtaining module 1301.

[0196] As an example of a software functional unit, the obtaining module 1301 can include code running on a computing instance. The computing instance can include at least one of a physical host (computing device), a virtual machine, a container. Further, the computing instance can be one or more. For example, the obtaining module 1301 can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers for running the code can be distributed in the same region, or in different regions. Further, the multiple hosts / virtual machines / containers for running the code can be distributed in the same availability zone (AZ), or in different AZs, each AZ including one data center or multiple data centers in close geographical proximity. Generally, one region can include multiple AZs.

[0197] Similarly, the multiple hosts / virtual machines / containers for running the code can be distributed in the same virtual private cloud (VPC), or in multiple VPCs. Generally, one VPC is set in one region, and communication between two VPCs in the same region, or between VPCs in different regions, needs to be set in each VPC to set a communication gateway, and the interconnection between VPCs is realized through the communication gateway.

[0198] As an example of a hardware functional unit, the obtaining module 1301 can include at least one computing device, such as a server, etc. Alternatively, the obtaining module 1301 can also be a device implemented by an application-specific integrated circuit (ASIC) or a programmable logic device (PLD), etc. The PLD can be implemented by a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0199] The plurality of computing devices included in the obtaining module 1301 can be distributed in the same region or in different regions. The plurality of computing devices included in the obtaining module 1301 can be distributed in the same AZ or in different AZs. Similarly, the plurality of computing devices included in the obtaining module 1301 can be distributed in the same VPC or in multiple VPCs. The plurality of computing devices can be any combination of servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0200] It should be noted that, in other embodiments, the obtaining module 1301 can be configured to perform any step of the simultaneous interpretation method, the translation module 1302 can be configured to perform any step of the simultaneous interpretation method, the generation module 1303 can be configured to perform any step of the simultaneous interpretation method, the output module 1304 can be configured to perform any step of the simultaneous interpretation method, the storage module 1305 can be configured to perform any step of the simultaneous interpretation method, the voiceprint recognition module 1306 can be configured to perform any step of the simultaneous interpretation method, the voice driving module 1307 can be configured to perform any step of the simultaneous interpretation method, and the receiving module 1308 can be configured to perform any step of the simultaneous interpretation method.

[0201] The steps implemented by the obtaining module 1301, the translation module 1302, the generation module 1303, the output module 1304, the storage module 1305, the voiceprint recognition module 1306, the voice driving module 1307, and the receiving module 1308 can be specified as needed, and the entire function of the simultaneous interpretation device can be implemented by the obtaining module 1301, the translation module 1302, the generation module 1303, the output module 1304, the storage module 1305, the voiceprint recognition module 1306, the voice driving module 1307, and the receiving module 1308 respectively implementing different steps in the simultaneous interpretation method.

[0202] The hardware device of the embodiment of the present application is illustrated below.

[0203] For example, Figure 14 is a structural schematic diagram of a computing device 1400 provided by an embodiment of the present application. As shown in Figure 14 the computing device 1400 includes a bus 1402, a processor 1404, a memory 1406, and a communication interface 1408. The processor 1404, the memory 1406, and the communication interface 1408 communicate through the bus 1402. The computing device 1400 can be a server or a terminal device. It should be understood that the number of processors and memories in the computing device 1400 is not limited by the embodiments of the present application.

[0204] The bus 1402 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, or the like. For ease of representation, Figure 14 Only one line is used to represent the bus, but this does not mean that there is only one bus or only one type of bus. The bus 1402 can include a path for transmitting information between various components (e.g., the memory 1406, the processor 1404, the communication interface 1408) of the computing device 1400.

[0205] The processor 1404 can include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP), or the like.

[0206] The memory 1406 can include a volatile memory (e.g., random access memory (RAM)), and can also include a non-volatile memory (e.g., read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD)).

[0207] The memory 1406 stores executable program code, and the processor 1404 executes the executable program code to respectively implement the functions of the aforementioned modules including but not limited to the acquisition module, the translation module, the generation module, and the output module, thereby implementing the above-mentioned simultaneous interpretation method, such as the method 100 or the method 300. That is, the memory 1406 stores instructions for executing the simultaneous interpretation method.

[0208] Alternatively, the memory 1406 stores executable program code, and the processor 1404 executes the executable program code to implement the functions of the aforementioned simultaneous interpretation device 1300, thereby implementing the simultaneous interpretation method. That is, the memory 1406 stores instructions for executing the simultaneous interpretation method.

[0209] The communication interface 1408 enables communication among the computing device 1400 and other devices or communication networks using, for example but not limited to, a transceiver module such as a network interface card, a transceiver, etc.

[0210] The embodiments of the present application also provide a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, for example, a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a notebook computer, or a smart phone.

[0211] For example, Figure 15 is a structural schematic diagram of a computing device cluster provided by an embodiment of the present application. As Figure 15 shown, the computing device cluster includes at least one computing device 1400. The memory 1406 in one or more computing devices 1400 in the computing device cluster can store the same instructions for performing the simultaneous interpretation method.

[0212] In some possible implementations, the memory 1406 in one or more computing devices 1400 in the computing device cluster can also respectively store partial instructions for performing the simultaneous interpretation method. In other words, the combination of one or more computing devices 1400 can collectively execute the instructions for performing the simultaneous interpretation method.

[0213] It should be noted that the memories 1406 in different computing devices 1400 in the computing device cluster can store different instructions for respectively performing part of the functions of the simultaneous interpretation apparatus. That is, the instructions stored in the memories 1406 in different computing devices 1400 can implement the functions of one or more of the aforementioned modules including but not limited to the obtaining module, the translation module, the generating module, and the output module.

[0214] In some possible implementations, one or more computing devices in the computing device cluster can be connected through a network. The network can be a wide area network or a local area network, etc. For example, Figure 16 is a structural schematic diagram of another computing device cluster provided by an embodiment of the present application. As Figure 16 shown, two computing devices 1400A and 1400B are connected through a network. Specifically, the communication interface in each computing device is connected to the network. In this type of possible implementation, the memory 1406 in the computing device 1400A stores instructions for performing the functions of the obtaining module and the output module. Meanwhile, the memory 1406 in the computing device 1400B stores instructions for performing the functions of the translation module and the generating module.

[0215] Figure 16The connection manner between the illustrated computing device clusters can be that the translation module and the generation module are implemented to perform the functions of the simultaneous interpretation method provided in the embodiments of the present application.

[0216] It should be understood that Figure 16 The functions of the computing device 1400A illustrated in the above can also be completed by multiple computing devices 1400. Similarly, the functions of the computing device 1400B can also be completed by multiple computing devices 1400.

[0217] The embodiments of the present application also provide another computing device cluster. The connection relationship between the computing devices in the computing device cluster can be similar to the connection manner of the computing device cluster illustrated in the above. Figure 15 and Figure 16 The connection manner of the computing device cluster illustrated in the above. The difference is that the memory 1406 in one or more computing devices 1400 in the computing device cluster can store the same instructions for executing the simultaneous interpretation method.

[0218] In some possible implementation manners, the memory 1406 in one or more computing devices 1400 in the computing device cluster can also respectively store partial instructions for executing the simultaneous interpretation method. In other words, the combination of one or more computing devices 1400 can collectively execute the instructions for executing the simultaneous interpretation method.

[0219] It should be noted that the memory 1406 in different computing devices 1400 in the computing device cluster can store different instructions for executing partial functions of the simultaneous interpretation apparatus. That is, the instructions stored in the memory 1406 in different computing devices 1400 can implement the functions of the simultaneous interpretation apparatus.

[0220] The embodiments of the present application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions, which can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device is caused to execute the simultaneous interpretation method.

[0221] The embodiments of the present application also provide a computer readable storage medium. The computer readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk), etc. The computer readable storage medium contains instructions, which instruct the computing device to execute the simultaneous interpretation method.

[0222] Finally, it should be noted that: the above examples are used to illustrate the technical solutions of the present application, but not limited to them; although the present application is described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A simultaneous interpretation method, characterized by, The method comprises: obtaining a character feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, the character feature comprising a pronunciation feature, and the first text being used to describe the voice content of the first voice; translating the first text into a second text, the first text being in a first language, and the second text being in a second language, the second language being different from the first language; generating a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, the voice content of the second voice being the description content of the second text, the pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and the emotion state reflected by the second voice being the same as the emotion state reflected by the first voice; outputting a simultaneous interpretation result corresponding to the first voice, the simultaneous interpretation result comprising the second voice.

2. The method of claim 1, wherein, The obtaining of the character feature corresponding to the first voice comprises: performing voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice; matching the target voiceprint feature with voiceprint features in a target feature set library, the target feature set library comprising a plurality of voiceprint feature-character feature correspondence relationships; if the target voiceprint feature is included in the target feature set library, taking the character feature corresponding to the target voiceprint feature in the target feature set library as the character feature corresponding to the first voice.

3. The method of claim 2, wherein, The method further comprises: if the target voiceprint feature is not included in the target feature set library, performing pronunciation feature extraction on the first voice to obtain a pronunciation feature corresponding to the first voice; storing the target voiceprint feature and the pronunciation feature corresponding to the first voice in the target feature set library correspondingly.

4. The method according to claim 2 or 3, characterized in that, The method further comprises: obtaining user registration information, the user registration information comprising a recorded voice; performing voiceprint recognition on the recorded voice to obtain a voiceprint feature corresponding to the recorded voice, and performing pronunciation feature extraction on the recorded voice to obtain a pronunciation feature corresponding to the recorded voice; storing the voiceprint feature corresponding to the recorded voice and the pronunciation feature corresponding to the recorded voice in the target feature set library correspondingly.

5. The method according to any one of claims 1 to 4, characterized in that, The character feature further comprises a character image, the character image comprising a face of a target character, and the method further comprises: driving the character image by the second voice to obtain a character video, the mouth movement of the target character in the character video being matched with the second voice, and the simultaneous interpretation result further comprising the character video.

6. The method according to any one of claims 1 to 5, characterized in that, The obtaining of the emotion category corresponding to the first voice comprises: performing emotion recognition on the first voice to obtain the emotion category corresponding to the first voice.

7. The method according to any one of claims 1 to 6, characterized in that, The generating of the second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text comprises: inputting the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text into a voice synthesis model to obtain the second voice output by the voice synthesis model.

8. The method according to any one of claims 1 to 7, characterized in that, The method is applied to a conference system, the first voice is from a first conference terminal in the conference system, and the outputting of the simultaneous interpretation result corresponding to the first voice comprises: sending the simultaneous interpretation result to a second conference terminal in the conference system.

9. The method of claim 8, wherein, The method further comprises: receiving a language selection instruction sent by the second conference terminal, the language selection instruction being used to indicate selection of the second language.

10. A simultaneous interpretation device, characterized by comprising: The device comprises: an obtaining module, configured to obtain a character feature corresponding to a first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, the character feature comprising a pronunciation feature, and the first text being used to describe voice content of the first voice; a translating module, configured to translate the first text into a second text, the first text being in a first language, and the second text being in a second language, the second language being different from the first language; a generating module, configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, voice content of the second voice being description content of the second text, the pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and an emotion state reflected by the second voice being the same as an emotion state reflected by the first voice; an outputting module, configured to output a simultaneous interpretation result corresponding to the first voice, the simultaneous interpretation result comprising the second voice.

11. The apparatus of claim 10, wherein, The obtaining module is configured to: perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice; match the target voiceprint feature with voiceprint features in a target feature set library, the target feature set library comprising a plurality of corresponding relationships between voiceprint features and character features; if the target voiceprint feature is included in the target feature set library, use a character feature corresponding to the target voiceprint feature in the target feature set library as the character feature corresponding to the first voice.

12. The apparatus of claim 11, wherein, The device further comprises a storage module. The obtaining module is further configured to, if the target voiceprint feature is not included in the target feature set library, perform pronunciation feature extraction on the first voice to obtain the pronunciation feature corresponding to the first voice. The storage module is configured to store the target voiceprint feature and the pronunciation feature corresponding to the first voice in the target feature set library correspondingly.

13. The apparatus of claim 11 or 12, wherein, The device further comprises a voiceprint recognition module and a storage module. The obtaining module is further configured to obtain user registration information, the user registration information comprising a recorded voice. The voiceprint recognition module is configured to perform voiceprint recognition on the recorded voice to obtain a voiceprint feature corresponding to the recorded voice, and perform pronunciation feature extraction on the recorded voice to obtain a pronunciation feature corresponding to the recorded voice. The storage module is configured to store the voiceprint feature corresponding to the recorded voice and the pronunciation feature corresponding to the recorded voice in the target feature set library correspondingly.

14. The apparatus of any one of claims 10 to 13, wherein, The character feature further comprises a character image, the character image comprising a face of a target character, and the device further comprises a voice driving module. The voice driving module is configured to drive the character image by using the second voice to obtain a character video, in which the mouth movement of the target character matches the second voice, and the simultaneous interpretation result further includes the character video.

15. The apparatus of any one of claims 10 to 14, wherein, The acquisition module is configured to: perform emotion recognition on the first voice to obtain an emotion category corresponding to the first voice.

16. The apparatus of any one of claims 10 to 15, wherein, The generation module is configured to: input the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text into a voice synthesis model to obtain the second voice output by the voice synthesis model.

17. The apparatus of any one of claims 10 to 16, wherein, The device is applied to a conference system, and the first voice is from a first conference terminal in the conference system; and the output module is configured to: send the simultaneous interpretation result to a second conference terminal in the conference system.

18. The apparatus of claim 17, wherein, The device further includes a receiving module. The receiving module is configured to receive a language selection instruction sent by the second conference terminal, the language selection instruction being used to indicate selection of the second language.

19. A conferencing system, characterized by The conference system includes a simultaneous interpretation server, a conference service platform, and a plurality of conference terminals, the plurality of conference terminals communicate through the conference service platform, the conference service platform is connected with the simultaneous interpretation server, and the plurality of conference terminals include a first conference terminal and a second conference terminal; The conference service platform is configured to receive a first voice sent by the first conference terminal and send the first voice to the simultaneous interpretation server; The simultaneous interpretation server is configured to acquire character features corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text, the character features include pronunciation features, the first text is used to describe the voice content of the first voice, the first text is in a first language, and the second text is in a second language, the second language being different from the first language; The simultaneous interpretation server is further configured to generate a second voice according to the pronunciation features corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform, the simultaneous interpretation result including the second voice, the voice content of the second voice being the description content of the second text, the pronunciation features corresponding to the second voice being consistent with the pronunciation features corresponding to the first voice, and the emotion state reflected by the second voice being the same as the emotion state reflected by the first voice; The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

20. A conferencing system, characterized by The conference system includes a simultaneous interpretation server, a conference service platform, a virtual interpreter terminal, and a plurality of conference terminals, the plurality of conference terminals communicate through the conference service platform, the virtual interpreter terminal is connected with the conference service platform and the simultaneous interpretation server respectively, and the plurality of conference terminals include a first conference terminal and a second conference terminal; The conference service platform is configured to receive the first voice sent by the first conference terminal and send the first voice to the virtual interpreter terminal; The virtual interpreter terminal is configured to send the first voice to the simultaneous interpretation server; The simultaneous interpretation server is configured to obtain a character feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text, the character feature including a pronunciation feature, the first text being used to describe voice content of the first voice, the first text being in a first language, and the second text being in a second language different from the first language; The simultaneous interpretation server is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the virtual interpreter terminal, the simultaneous interpretation result including the second voice, voice content of the second voice being description content of the second text, the pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and an emotion state reflected by the second voice being the same as an emotion state reflected by the first voice; The virtual interpreter terminal is configured to send the simultaneous interpretation result to the conference service platform; The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

21. The conferencing system of claim 19 or 20, wherein, The simultaneous interpretation server has a plurality of simultaneous interpretation engines deployed therein, wherein each of the simultaneous interpretation engines corresponds to one conference terminal in the conference system, or each of the simultaneous interpretation engines corresponds to one conference held through the conference system, or each of the simultaneous interpretation engines corresponds to one language, the language being a language to be translated or a target language to be translated.

22. A conferencing system, characterized by The conference system includes a conference service platform and a plurality of conference terminals, the plurality of conference terminals being in communication through the conference service platform, and the plurality of conference terminals including a first conference terminal and a second conference terminal; The conference service platform is configured to receive a first voice sent by the first conference terminal, obtain a character feature corresponding to the first voice, and send the first voice and the character feature corresponding to the first voice to the second conference terminal, the character feature including a pronunciation feature; The second conference terminal is configured to obtain an emotion category corresponding to the first voice and a first text corresponding to the first voice, and translate the first text into a second text, the first text being used to describe voice content of the first voice, the first text being in a first language, and the second text being in a second language different from the first language; The second conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and play a simultaneous interpretation result corresponding to the first voice, the simultaneous interpretation result comprising the second voice, a voice content of the second voice being description content of the second text, the pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and an emotion state reflected by the second voice being the same as an emotion state reflected by the first voice.

23. The conferencing system of claim 22, wherein, The conference system further comprises a simultaneous interpretation server connected with the conference service platform, and a target feature set library is stored in the simultaneous interpretation server, the target feature set library comprising a plurality of corresponding relationships between voiceprint features and person features; The conference service platform is configured to perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice, and send the target voiceprint feature to the simultaneous interpretation server; The simultaneous interpretation server is configured to match the target voiceprint feature with voiceprint features in the target feature set library, and if the target feature set library comprises the target voiceprint feature, send a person feature corresponding to the target voiceprint feature in the target feature set library to the conference service platform; The conference service platform is configured to take the person feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the person feature corresponding to the first voice.

24. A conferencing system, characterized by The conference system comprises a conference service platform and a plurality of conference terminals, the plurality of conference terminals being in communication through the conference service platform, and the plurality of conference terminals comprising a first conference terminal and a second conference terminal; The conference service platform is configured to receive a first voice sent by the first conference terminal, and send the first voice to the second conference terminal; The second conference terminal is configured to obtain a person feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text, the person feature comprising a pronunciation feature, the first text being used to describe a voice content of the first voice, the first text being in a first language, and the second text being in a second language different from the first language; The second conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and play a simultaneous interpretation result corresponding to the first voice, the simultaneous interpretation result comprising the second voice, a voice content of the second voice being description content of the second text, the pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and an emotion state reflected by the second voice being the same as an emotion state reflected by the first voice.

25. The conferencing system of claim 24, wherein, The conference system further comprises a simultaneous interpretation server, the plurality of conference terminals are connected with the simultaneous interpretation server respectively, the simultaneous interpretation server stores a target feature set library, and the target feature set library comprises a plurality of corresponding relationships between voiceprint features and character features; The second conference terminal is configured to perform voiceprint recognition on the first voice to obtain target voiceprint features corresponding to the first voice, and send the target voiceprint features to the simultaneous interpretation server; The simultaneous interpretation server is configured to match the target voiceprint features with voiceprint features in the target feature set library, and if the target voiceprint features are included in the target feature set library, send character features corresponding to the target voiceprint features in the target feature set library to the second conference terminal; The second conference terminal is configured to take the character features corresponding to the target voiceprint features sent by the simultaneous interpretation server as character features corresponding to the first voice.

26. A conferencing system, characterized by The conference system comprises a conference service platform and a plurality of conference terminals, the plurality of conference terminals communicate through the conference service platform, and the plurality of conference terminals comprise a first conference terminal and a second conference terminal; The first conference terminal is configured to collect a first voice and send the first voice to the conference service platform; The conference service platform is configured to obtain character features corresponding to the first voice, and send the character features corresponding to the first voice to the second conference terminal, wherein the character features comprise pronunciation features; The first conference terminal is configured to obtain an emotional category corresponding to the first voice and a first text corresponding to the first voice, and translate the first text into a second text, wherein the first text is used to describe voice content of the first voice, the first text is in a first language, and the second text is in a second language different from the first language; The first conference terminal is further configured to generate a second voice according to the pronunciation features corresponding to the first voice, the emotional category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform, wherein the simultaneous interpretation result comprises the second voice, voice content of the second voice is description content of the second text, pronunciation features of the second voice are consistent with the pronunciation features corresponding to the first voice, and an emotional state reflected by the second voice is the same as an emotional state reflected by the first voice; The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

27. The conferencing system of claim 26, wherein, The conference system further comprises a simultaneous interpretation server, the conference service platform is connected with the simultaneous interpretation server, the simultaneous interpretation server stores a target feature set library, and the target feature set library comprises a plurality of corresponding relationships between voiceprint features and character features; The conference service platform is configured to perform voiceprint recognition on the first voice to obtain target voiceprint features corresponding to the first voice, and send the target voiceprint features to the simultaneous interpretation server; The simultaneous interpretation server is configured to match the target voiceprint feature with voiceprint features in the target feature set library, and if the target voiceprint feature is included in the target feature set library, send a person feature corresponding to the target voiceprint feature in the target feature set library to the conference service platform. The conference service platform is configured to take the person feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the person feature corresponding to the first voice.

28. A conferencing system, characterized by The conference system comprises a conference service platform and a plurality of conference terminals, the plurality of conference terminals communicate through the conference service platform, and the plurality of conference terminals comprise a first conference terminal and a second conference terminal. The first conference terminal is configured to collect a first voice, obtain a person feature corresponding to the first voice, an emotion category corresponding to the first voice, and a first text corresponding to the first voice, and translate the first text into a second text, the person feature comprises a pronunciation feature, the first text is used to describe the voice content of the first voice, the first text is in a first language, and the second text is in a second language, the second language being different from the first language. The first conference terminal is further configured to generate a second voice according to the pronunciation feature corresponding to the first voice, the emotion category corresponding to the first voice, and the second text, and send a simultaneous interpretation result corresponding to the first voice to the conference service platform, the simultaneous interpretation result comprising the second voice, the voice content of the second voice being the description content of the second text, the pronunciation feature corresponding to the second voice being consistent with the pronunciation feature corresponding to the first voice, and the emotion state reflected by the second voice being the same as the emotion state reflected by the first voice. The conference service platform is configured to send the simultaneous interpretation result to the second conference terminal.

29. The conferencing system of claim 28, wherein, The conference system further comprises a simultaneous interpretation server, the first conference terminal is connected with the simultaneous interpretation server, and the simultaneous interpretation server stores a target feature set library, the target feature set library comprises a plurality of corresponding relationships between voiceprint features and person features. The first conference terminal is configured to perform voiceprint recognition on the first voice to obtain a target voiceprint feature corresponding to the first voice, and send the target voiceprint feature to the simultaneous interpretation server. The simultaneous interpretation server is configured to match the target voiceprint feature with voiceprint features in the target feature set library, and if the target voiceprint feature is included in the target feature set library, send a person feature corresponding to the target voiceprint feature in the target feature set library to the first conference terminal. The first conference terminal is configured to take the person feature corresponding to the target voiceprint feature sent by the simultaneous interpretation server as the person feature corresponding to the first voice.

30. The conference system of any of claims 19 to 29, wherein The first conference terminal is configured to display a conference registration interface, and the conference registration interface displays a person information input option, the person information input option being used for a user to provide person information, and the person information comprising a recorded voice. The first conference terminal is further configured to generate user registration information in response to receiving the recorded voice through the conference registration interface, the user registration information comprising the recorded voice, the recorded voice being used to determine a correspondence between a set of voiceprint features and pronunciation features.

31. The conference system of any of claims 19 to 30, wherein, The second conference terminal is configured to receive a language selection instruction through the display interface, the language selection instruction being used to indicate selection of the second language.

32. A cluster of computing devices, characterized in that, at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method of any of claims 1 to 9.

33. A computer program product comprising instructions, wherein: The instructions, when executed by the cluster of computing devices, cause the cluster of computing devices to perform the method of any of claims 1 to 9.

34. A computer-readable storage medium, characterized in that, computer program instructions, the cluster of computing devices performing the method of any of claims 1 to 9 when the computer program instructions are executed by the cluster of computing devices.