A translation progress display method and related device

By displaying the interpretation progress on the speaker's device, the problem of simultaneous interpretation delay is solved, helping speakers adjust their pace in a timely manner, ensuring that the translated content is synchronized with the spoken content, and improving the meeting effect.

CN122157663APending Publication Date: 2026-06-05HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2024-12-03
Publication Date
2026-06-05

Smart Images

  • Figure CN122157663A_ABST
    Figure CN122157663A_ABST
Patent Text Reader

Abstract

The application discloses a kind of translation progress display method, applied to simultaneous interpretation process.In the translation progress display method, by displaying the translation progress of the speech content of the speaker on the electronic device used by the speaker, the speaker can understand the current dissemination progress of the translation content in time, so as to adaptively control the speaking pace of himself, avoid the translation content heard by the audience far behind the actual speech content of the speaker, and ensure the conference communication effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of translation technology, and in particular to a method and related apparatus for displaying translation progress. Background Technology

[0002] In multilingual conference settings, simultaneous interpretation is often necessary to enable all participants to communicate across languages. This means translating the speaker's remarks to the audience without interruption.

[0003] Currently, in most business scenarios, simultaneous interpretation is primarily handled by conference system software. Specifically, speakers and audience members join the same meeting using their respective electronic devices. The speaker's device sends their raw speech data to a server equipped with the conference system software. The server then translates the raw speech data, producing translated speech data. This translated speech data is then sent to the audience members' electronic devices, thus completing the simultaneous interpretation process.

[0004] However, current simultaneous interpretation processes have a certain interpretation delay, which can easily cause the translated content heard by the audience to lag far behind the actual content of the speaker's speech, thus affecting the effectiveness of conference communication. Summary of the Invention

[0005] This application provides a method for displaying the translation progress, enabling speakers to understand the current dissemination progress of the translated content in a timely manner, thereby adaptively controlling their speaking pace.

[0006] Firstly, a method for displaying interpretation progress is provided, applied to a first electronic device used by a speaker. This method includes: the first electronic device acquiring a first interpretation progress. The first interpretation progress indicates the propagation progress of a first translated content, which is the content obtained by translating the original speech data captured by the first electronic device. The language corresponding to the first translated content is different from the language corresponding to the original speech data. For example, the language corresponding to the original speech data is Chinese, while the language corresponding to the first translated content is English, Japanese, Korean, or other languages.

[0007] Then, the first electronic device displays the first interpretation progress so that the speaker can keep track of the current dissemination progress of the first translated content by observing the first interpretation progress.

[0008] In this solution, the translation progress of the speaker's speech is displayed on the electronic device used by the speaker, allowing the speaker to understand the current dissemination progress of the translated content in a timely manner. This enables the speaker to adapt and control their speaking pace, preventing the translated content heard by the audience from lagging far behind the speaker's actual speech, and ensuring the effectiveness of the meeting communication.

[0009] In one possible implementation, the first electronic device displays the first translation progress, specifically including: the first electronic device displays the translated content and / or the untranslated content in the first translated content.

[0010] In this solution, by displaying the disseminated and / or undisseminated content of the first translation on the first electronic device, the speaker can quickly understand how much content is still to be translated, thereby clarifying how long it will take for the current translation progress to catch up with their actual speaking progress, and thus adaptively adjusting their speaking pace.

[0011] In one possible implementation, the first electronic device displays a first interpretation progress, specifically including: the first electronic device displays a progress visualization indicator, which is used to indicate the first interpretation progress. The progress visualization indicator is, for example, a combination of one or more of a pattern, text, and color.

[0012] In one possible implementation, progress visualization markers are used to indicate the progress of sentence propagation and / or word propagation of the first translated content.

[0013] The progress of sentence propagation can be achieved by displaying the number of unpropagated sentences, the ratio between the number of propagated and unpropagated sentences, the ratio between the length of propagated and unpropagated sentences, or by displaying the information of sentences currently being propagated. Similarly, the progress of word propagation can be achieved by displaying the number of unpropagated words or the ratio between the number of propagated and unpropagated words.

[0014] In this solution, the progress of statements and / or words is indicated by displaying progress visualization indicators. Speakers can clearly know how many statements and / or words have not yet been propagated, and thus adapt their speaking pace accordingly.

[0015] In one possible implementation, a progress visualization indicator is used to indicate the expected propagation time of the unpropagated content in the first translation, that is, how much longer it will take for the unpropagated content to be propagated.

[0016] In one possible implementation, to further assist the speaker in controlling their speaking pace, when the untranslated content in the first translated content meets preset conditions, the first electronic device displays a prompt message to remind the speaker to slow down their speaking speed.

[0017] In one possible implementation, the preset conditions include that the number of statements included in the unspread content is greater than or equal to the target threshold, or that the spread duration of the unspread content is greater than or equal to the target duration.

[0018] In one possible implementation, the first translation content is translated text, and the language of the translated text is different from the language of the original speech data; or, the first translation content is translated speech data, and the language of the translated speech data is different from the language of the original speech data.

[0019] In one possible implementation, the first electronic device acquires a second translation progress, which is used to indicate the progress of transmitting a second translated content, the second translated content being the content obtained by translating the original speech data, and the language corresponding to the second translated content is different from the language corresponding to the first translated content; the first electronic device displays the second translation progress.

[0020] That is, the first electronic device used by the speaker can display the interpretation progress corresponding to different translated content, so that the speaker can clearly know the interpretation progress corresponding to the translated content in various languages.

[0021] In one possible implementation, the first electronic device displays a second translation progress, including: the first electronic device simultaneously displays a first translation progress and a second translation progress; or, in response to receiving a display switching instruction, the first electronic device switches from displaying the first translation progress to displaying the second translation progress.

[0022] In one possible implementation, the first electronic device acquires the voiceprint of the original speech data; then, the first electronic device displays the voiceprint of the original speech data.

[0023] In this scheme, by displaying the voiceprint of the original speech data, the speaker can understand the rhythm of the previous speech process (such as changes in speaking volume, speaking speed, and pauses), which helps the speaker to adaptively adjust the rhythm of subsequent speeches.

[0024] In one possible implementation, a first electronic device is located in a conference system, which translates the raw voice data into first translated content and transmits the first translated content to a second electronic device in the conference system, which plays the first translated content via voice.

[0025] In one possible implementation, the first electronic device obtains the first translation progress, specifically including: the first electronic device sending raw voice data to the server; the first electronic device receiving the first translated content and / or the expected propagation duration of the first translated content sent by the server; and the first electronic device determining the first translation progress based on the first translated content and / or the expected propagation duration of the first translated content.

[0026] In a second aspect, a method for displaying translation progress is provided, comprising: a server receiving raw voice data sent by a first electronic device; the server performing translation on the raw voice data to obtain first translated content; the server sending the first translated content and / or the expected propagation duration of the first translated content to the first electronic device, wherein the first translated content and / or the expected propagation duration of the first translated content are used to enable the first electronic device to display a first translation progress, wherein the first translation progress is used to indicate the progress of propagating the first translated content.

[0027] In one possible implementation, the server sends first translated content and digital human video to a second electronic device; wherein the digital human video is used to demonstrate the digital human's speech process, and the changes in the digital human's mouth shape during the speech process are determined based on the first translated content.

[0028] Thirdly, an interpretation progress display device is provided, which is deployed on a first electronic device. The interpretation progress display device includes: an acquisition module for acquiring a first interpretation progress, which indicates the propagation progress of a first translated content, the first translated content being the content obtained by translating the original speech data captured by the first electronic device; and a display module for displaying the first interpretation progress.

[0029] In one possible implementation, the display module is specifically used to: display the propagated and / or unpropagated content in the first translated content.

[0030] In one possible implementation, the display module is specifically used to: display a progress visualization indicator, which is used to indicate the progress of the first interpretation.

[0031] In one possible implementation, progress visualization markers are used to indicate the progress of sentence propagation and / or word propagation of the first translated content.

[0032] In one possible implementation, progress visualization markers are used to indicate the expected propagation time of unpropagated content in the first translation.

[0033] In one possible implementation, the display module is further configured to: display a prompt message when the untranslated content in the first translated content meets a preset condition, the prompt message being used to prompt the speaker to slow down their speaking speed.

[0034] In one possible implementation, the preset conditions include that the number of statements included in the unspread content is greater than or equal to the target threshold, or that the spread duration of the unspread content is greater than or equal to the target duration.

[0035] In one possible implementation, the first translation content is translated text, and the language of the translated text is different from the language of the original speech data; or, the first translation content is translated speech data, and the language of the translated speech data is different from the language of the original speech data.

[0036] In one possible implementation, the acquisition module is further configured to acquire a second translation progress, which indicates the progress of transmitting the second translated content. The second translated content is the content obtained by translating the original speech data, and the language corresponding to the second translated content is different from the language corresponding to the first translated content. The display module is further configured to display the second translation progress.

[0037] In one possible implementation, the display module is specifically configured to: simultaneously display the first interpretation progress and the second interpretation progress; or, in response to receiving a display switching instruction, switch from displaying the first interpretation progress to displaying the second interpretation progress.

[0038] In one possible implementation, the acquisition module is also used to acquire the voiceprint of the raw speech data; the display module is also used to display the voiceprint.

[0039] In one possible implementation, a first electronic device is located in a conference system, which translates the raw voice data into first translated content and transmits the first translated content to a second electronic device in the conference system, which plays the first translated content via voice.

[0040] In one possible implementation, the acquisition module is specifically used for: sending raw voice data to the server; receiving first translated content and / or the expected propagation duration of the first translated content sent by the server; and determining the first translation progress based on the first translated content and / or the expected propagation duration of the first translated content.

[0041] Fourthly, an interpretation progress display device is provided, which is deployed on a server. The interpretation progress display device includes: a transceiver module for receiving raw voice data sent by a first electronic device; a processing module for translating the raw voice data to obtain first translated content; the transceiver module is further used to send the first translated content and / or the expected propagation duration of the first translated content to the first electronic device, the first translated content and / or the expected propagation duration of the first translated content being used to enable the first electronic device to display a first interpretation progress, the first interpretation progress being used to indicate the progress of propagating the first translated content.

[0042] In one possible implementation, the transceiver module is further configured to: send the first translated content and the digital human video to the second electronic device; wherein the digital human video is used to demonstrate the digital human's speaking process, and the changes in the digital human's mouth shape during the speaking process are determined based on the first translated content.

[0043] Fifthly, an electronic device is provided, which may include a processor and a memory coupled together. The memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method described in the first aspect or any implementation thereof is implemented. For details regarding the steps of the various possible implementations of the first aspect executed by the processor, please refer to the first aspect; further details will not be elaborated here.

[0044] Sixthly, a server is provided, which may include a processor and a memory coupled together. The memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method described in the second aspect or any implementation thereof is implemented. For details regarding the steps of the various possible implementations of the first aspect executed by the processor, please refer to the first aspect; they will not be repeated here.

[0045] In a seventh aspect, a conference system is provided, comprising the electronic device described in the fifth aspect and the server described in the sixth aspect.

[0046] Eighthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, which, when run on a computer, causes the computer to perform the method of any implementation of the first aspect described above.

[0047] Ninth aspect, a circuit system is provided, the circuit system including processing circuitry, the processing circuitry being configured to perform the method of any implementation of the first aspect described above.

[0048] In a tenth aspect, a computer program product is provided that, when run on a computer, causes the computer to execute any implementation of the first aspect described above.

[0049] Eleventhly, a chip system is provided, comprising a processor for supporting a server or threshold acquisition device in implementing the functions involved in any implementation of the first aspect, such as transmitting or processing data and / or information involved in the above methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for the server or communication device. The chip system may be composed of chips or may include chips and other discrete devices.

[0050] The beneficial effects of the second to eleventh aspects mentioned above can be referred to the introduction of the first aspect above, and will not be repeated here. Attached Figure Description

[0051] Figure 1 A schematic diagram of a system architecture provided for this application;

[0052] Figure 2 A flowchart illustrating an interpretation progress display method provided in this application;

[0053] Figure 3 This application provides a schematic diagram showing disseminated and undisseminated content;

[0054] Figure 4 This application provides a schematic diagram of using a pie chart to represent the progress of interpretation.

[0055] Figure 5 This application provides a schematic diagram for representing the translation progress using text.

[0056] Figure 6 This application provides a schematic diagram of using a progress bar to represent the interpretation progress;

[0057] Figure 7 This application provides a schematic diagram of a display of prompt information;

[0058] Figure 8 A schematic diagram illustrating a switching display of interpretation progress provided in this application;

[0059] Figure 9 A schematic diagram illustrating the display of a voiceprint provided in this application;

[0060] Figure 10 A flowchart illustrating a simultaneous interpretation process in a conference, provided for the purposes of this application;

[0061] Figure 11 This application provides a schematic diagram illustrating the conversion of translated text into translated speech data.

[0062] Figure 12 A schematic diagram of an interpretation progress display device provided in this application;

[0063] Figure 13 A schematic diagram of another interpretation progress display device provided in this application;

[0064] Figure 14 A schematic diagram of the execution device provided in this application;

[0065] Figure 15 This is a schematic diagram of the structure of a computer-readable storage medium provided in this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application are described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some, and not all, of the embodiments of this application. Those skilled in the art will understand that, with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0067] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such descriptions can be used interchangeably where appropriate to allow embodiments to be implemented in a sequence other than that illustrated or described in this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to such processes, methods, products, or devices. The naming or numbering of steps appearing in this application does not imply that the steps in the method flow must be performed in the chronological / logical order indicated by the naming or numbering. The execution order of named or numbered process steps can be changed according to the desired technical purpose, as long as the same or similar technical effect is achieved.

[0068] The division of units in this application is a logical division. In practical applications, there may be other division methods. For example, multiple units may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the shown or discussed mutual coupling, direct coupling, or communication connection may be through some interface, and the indirect coupling or communication connection between units may be electrical or other similar forms, none of which are limited in this application. Furthermore, the units or sub-units described as separate components may or may not be physically separated, may or may not be physical units, or may be distributed among multiple circuit units. Some or all of the units can be selected to achieve the purpose of the solution in this application according to actual needs.

[0069] Currently, simultaneous interpretation processes often involve a certain interpretation delay, which can cause the translated content heard by the audience to lag far behind the actual content spoken by the speaker, thus affecting the effectiveness of conference communication.

[0070] The applicant's research revealed two main factors contributing to the time delay in simultaneous interpreting: 1. The inherent lag in translation itself, meaning that translation can only begin after the speaker has finished speaking a complete sentence. 2. Differences in expression habits between languages; a small portion of information expressed in one language may require a much longer portion to convey clearly in another. In other words, the translated speech may be longer, resulting in a longer playback time for the translated audio data.

[0071] In summary, the combination of the two factors mentioned above can lead to significant time delays in simultaneous interpreting. Furthermore, if the speaker is unaware of this existing delay and continues speaking at a rapid pace, the translated content may lag far behind the speaker's actual delivery, thus impacting the effectiveness of the meeting.

[0072] For example, in some meeting scenarios, speakers need to combine their presentations with the content of the slides, and the content of their speech is closely related to the content of the slides. In this case, if the translated content lags far behind the speaker's actual speech, it may result in a situation where the speaker has already turned the page of the slide, but the audience is still hearing the translated content based on the previous page of the slide (i.e., audio-visual desynchronization).

[0073] For example, in interactive conference settings, participants need to communicate frequently, and audience members may raise questions about the speaker's remarks at any time. If the translated content lags far behind the speaker's actual remarks, the translated content that the audience is asking may be significantly out of sync with the speaker's current remarks, making effective communication between the audience and the speaker difficult.

[0074] In view of this, this application provides a method for displaying interpretation progress. By displaying the interpretation progress of the speaker's speech on the electronic device used by the speaker, the speaker can understand the current dissemination progress of the translated content in a timely manner, thereby adaptively controlling their own speaking pace and preventing the translated content heard by the audience from lagging far behind the speaker's actual speech content, thus ensuring the effectiveness of the meeting communication.

[0075] Please see Figure 1 , Figure 1 This is a schematic diagram of a system architecture provided for this application. (For example...) Figure 1 As shown, the system architecture includes server 001 and one or more electronic devices 002. Optionally, server 001 may specifically include one or more physical servers.

[0076] Electronic device 002 can join an online meeting by installing and launching conferencing software, thereby establishing a connection with server 001. Through server 001, it can share audio, video, and display content with other electronic devices joining the same meeting. Server 001 provides simultaneous interpretation services by receiving raw audio data from the speaker's electronic device 002, translating it into audio data in other languages, and then sending the translated audio data to the audience's electronic devices 002, thus completing the simultaneous interpretation process.

[0077] Furthermore, in this application, the server 001 can also send the translated text content or voice data to the electronic device 002 used by the speaker, so that the electronic device 002 used by the speaker can display the translation progress, so that the speaker can control the speaking pace.

[0078] Among them, electronic device 002 can be, for example, a smartphone, a personal computer (PC), a laptop, a tablet computer, a smart TV, a mobile internet device (MID), a display device in an autonomous vehicle, a display device in a smart home, and other such devices.

[0079] Please see Figure 2 , Figure 2 This is a flowchart illustrating a method for displaying interpretation progress provided in this application. Figure 2 As shown, the translation progress display method provided in this application includes the following steps 201-202.

[0080] Step 201: The first electronic device acquires the first translation progress. The first translation progress is used to indicate the propagation progress of the first translated content. The first translated content is the content obtained by translating the original speech data captured by the first electronic device.

[0081] In this application, the first electronic device is an electronic device used by the speaker, and therefore, the first electronic device can capture the speaker's original voice data through a microphone or other device. Furthermore, the original voice data captured by the first electronic device is translated into first translated content and transmitted to the electronic devices used by the listeners. The language corresponding to the first translated content is different from the language corresponding to the original voice data. For example, the language corresponding to the original voice data is Chinese, while the language corresponding to the first translated content is English, Japanese, Korean, or other languages.

[0082] Specifically, after the raw speech data captured by the first electronic device is translated into first translated content, the first translated content needs to be transmitted on the electronic device used by the listener to complete the simultaneous interpretation process. That is, the first translated content is actually played in the form of speech on the electronic device used by the listener. Due to the time delay in the simultaneous interpretation process, after the first electronic device captures the raw speech data, the first translated content corresponding to the raw speech data still needs some time to be transmitted. At this time, the first electronic device can obtain the first interpretation progress, which is the transmission progress of the first translated content.

[0083] Optionally, the first translation content is translated text, and the language of the translated text is different from the language of the original speech data. Alternatively, the first translation content is translated speech data, and the language of the translated speech data is different from the language of the original speech data.

[0084] When the first translated content is translated text, the progress of propagating the first translated content can be understood as the progress of playing the corresponding audio data. When the first translated content is translated audio data, the progress of propagating the first translated content can be understood as the progress of playing the translated audio data.

[0085] In this application, the first electronic device is, for example, located in a conference system. This conference system is used to translate the raw voice data into first translated content and transmit the first translated content to a second electronic device within the conference system. The second electronic device is used to play the first translated content via voice. Of course, the translation progress display method provided in this application can also be applied to other multi-device collaborative communication scenarios similar to conference settings, and this application does not specifically limit its application in this regard.

[0086] That is, the interpretation progress display method provided in this application is applied to simultaneous interpretation processes in conference scenarios. A first electronic device and a second electronic device are added to the same conference system, and the raw speech data captured by the first electronic device is translated into translated content in other languages, so that the second electronic device can play the translated content by voice to complete the simultaneous interpretation process.

[0087] Of course, in some embodiments, the conferencing system may also send the first translated content obtained from translating the original voice data to the first electronic device, so that the first electronic device can play the first translated content by voice or perform other processing operations based on the first translated content.

[0088] Step 202: The first electronic device displays the first translation progress.

[0089] Based on the acquired first interpretation progress, the first electronic device can display the first interpretation progress on a screen, allowing the speaker to understand the current dissemination progress of the first translated content and thus adaptively control their speaking pace. For example, if the first interpretation progress indicates that the current dissemination progress of the first translated content is slow, the speaker can pause slightly or slow down their speaking speed to avoid the audience hearing the translated content lagging far behind the speaker's actual speech.

[0090] It should be noted that, because the speaker may continue to speak and the translation is also continuously being translated during the interpretation process, the first interpretation progress displayed by the first electronic device often changes dynamically over time.

[0091] Optionally, there are several ways for the first electronic device to display the first interpretation progress.

[0092] Method 1: The first electronic device displays the disseminated and / or undisseminated content from the first translated content.

[0093] In other words, the first electronic device may display only the disseminated content of the first translated content, or only the undisseminated content of the first translated content. Alternatively, the first electronic device may display both the disseminated and undisseminated content of the first translated content simultaneously. When the first electronic device displays only the disseminated content of the first translated content, the speaker can easily determine, based on their actual speech, what content remains undisseminated (i.e., identify the undisseminated content).

[0094] Specifically, by displaying the transmitted and / or untransmitted content of the first translation on the first electronic device, the speaker can quickly understand how much content is still to be translated, thereby clarifying how long it will take for the current translation progress to catch up with their actual speaking progress, and thus adaptively adjusting their speaking pace.

[0095] It should be noted that when the first electronic device simultaneously displays both the disseminated and undisseminated content from the first translation, the first electronic device can employ different display methods for the disseminated and undisseminated content, thereby enabling the speaker to distinguish which content is disseminated and which is undisseminated. For example, the disseminated and undisseminated content can be rendered using different colors; or, the disseminated and undisseminated content can be represented using different font formats or font sizes.

[0096] Alternatively, the first electronic device could also mark the transmitted and untransmitted content separately, such as boxing the transmitted content and marking it as "transmitted" and boxing the untransmitted content and marking it as "untransmitted", so that the speaker can distinguish between the transmitted and untransmitted content.

[0097] Furthermore, when the first translation content includes a large amount of information, for ease of display, the first electronic device may only display the propagated and unpropagated content of the translated sentence currently being propagated within the first translation content. For example, assuming the first translation content includes multiple translated sentences, and one of these multiple translated sentences is currently being propagated, then the first electronic device may display the propagated and unpropagated content of this propagated translated sentence.

[0098] For example, please refer to Figure 3 , Figure 3 This application provides a schematic diagram showing disseminated and undisseminated content. For example... Figure 3 As shown, the electronic device used by the speaker displays the interface of a meeting that the device has joined. At the bottom of the meeting interface, there is an interpretation card that includes an interpretation progress bar. This progress bar displays the translation progress of the content in a specific language. Furthermore, in other implementations, the interpretation card can be placed in other locations on the meeting interface, such as the left or right side; this application does not specifically limit the location of the interpretation card.

[0099] For example, suppose the original speech data corresponds to Chinese speech, and the content of the original speech data is "This morning, I noticed an interesting phenomenon while taking a walk in the park." Then, in... Figure 3The translation progress bar shows the progress of the translation of the English version of the original audio data. Specifically, the English version of the original audio data is "This morning, while taking a walk in the park, I discovered an interesting phenomenon." Furthermore, "This morning" in the English version is the content that has already been broadcast and is displayed in bold and italic font; "while taking a walk in the park, I discovered an interesting phenomenon" is the content that has not yet been broadcast and is not displayed in bold or italic font. That is, "This morning" is content that has already been played on the electronic devices used by the listeners, while "while taking a walk in the park, I discovered an interesting phenomenon" is content that has not yet been played on the electronic devices used by the listeners.

[0100] Method 2: The first electronic device displays the sentence propagation progress and / or word propagation progress of the first translated content.

[0101] Specifically, the progress of sentence propagation can be achieved by displaying the number of unpropagated sentences, or by displaying the ratio between the number of propagated and unpropagated sentences, or by displaying the propagation duration of unpropagated sentences, or by displaying the propagation duration of propagated and unpropagated sentences, or by displaying the sentence information currently being propagated. Similarly, the progress of word propagation can be achieved by displaying the number of unpropagated words, or by displaying the ratio between the number of propagated and unpropagated words.

[0102] Among them, the number of unpropagated sentences is the number of sentences included in the unpropagated content of the first translation, and the number of unpropagated words is the number of words included in the unpropagated content.

[0103] Understandably, if the first translated content is in a language the speaker is unfamiliar with, the speaker may have difficulty understanding it if it is displayed as a translated text. Furthermore, if the first translated content needs to be displayed as a lengthy translated text, the space allocated for interpreting progress on the meeting interface may not be sufficient to display the entire first translated content.

[0104] Therefore, in method 2, the first electronic device may only display the progress of sentence propagation, or only the progress of word propagation. Alternatively, the first electronic device may display both the progress of sentence propagation and the progress of word propagation. In this way, based on the progress of sentence propagation and / or word propagation displayed by the first electronic device, the speaker can clearly know how many sentences and / or words have not yet been propagated, and thus adaptively adjust their speaking pace.

[0105] Specifically, the first electronic device may display the sentence propagation progress and / or word propagation progress of the first translated content through progress visualization indicators.

[0106] The progress visualization indicator can be a graphic symbol. For example, when the first electronic device indicates the progress of statement propagation by displaying the number of unpropagated statements, it can display one or more preset patterns (such as vertical bars, dots, or triangles), thereby representing the number of unpropagated statements based on the number of preset patterns. Alternatively, when the first electronic device indicates the progress of statement propagation by displaying the ratio between the number of propagated statements and the number of unpropagated statements, it can display a pie chart or bar chart to represent the ratio. Yet another example is that the first electronic device can display a pie chart or bar chart to represent the ratio between the length of propagated statements and the length of unpropagated statements.

[0107] For example, please refer to Figure 4 , Figure 4 This application provides a schematic diagram using a pie chart to represent the progress of interpretation. For example... Figure 4 As shown, in the interpretation card, the visual identifier can specifically be a pie chart. The black portion of the pie chart represents the number of statements that have been propagated, while the white portion represents the number of statements that have not been propagated. Therefore, this pie chart can be used to indicate the ratio between the number of propagated statements and the number of unpropagated statements. Of course, in some embodiments, the pie chart can also be used to indicate the ratio between the propagation time of propagated statements and the propagation time of unpropagated statements.

[0108] Furthermore, the progress visualization can also be text-based. That is, the first electronic device can also display the progress of sentence propagation and / or word propagation of the first translated content through text. For example, when the first electronic device indicates the progress of sentence propagation by displaying the number of unpropagated sentences, it could display the text "2 sentences remaining." As another example, when the first electronic device indicates the progress of word propagation by displaying the number of unpropagated words, it could display the text "12 words remaining." Yet another example, when the first electronic device indicates the progress of sentence propagation by displaying the information of the sentences currently being propagated, it could display the text "while taking a walk in the park," which represents the currently propagated sentence information.

[0109] For example, please refer to Figure 5 , Figure 5 This application provides a schematic diagram illustrating the translation progress using text. For example... Figure 5 As shown, in the interpretation card, the visual identifier can be a text that reads "2 sentences remaining unplayed," indicating that the number of unplayed sentences is 2.

[0110] In general, this application does not limit the specific implementation of the progress visualization mark, which can be composed of one or more of the following: pattern, text, and color.

[0111] Method 3: The first electronic device displays the expected propagation duration of the unpropagated content in the first translated content.

[0112] Specifically, the first electronic device can also display the expected propagation time of the unpropagated content in the first translated content through a progress visualization indicator.

[0113] For example, the progress visualization can be represented by a progress bar. The length of the progress bar indicates the expected propagation time of the unpropagated content in the first translated content; alternatively, the progress bar can be divided into two parts, with the first part indicating the propagation time of the propagated content in the first translated content, and the second part indicating the expected propagation time of the unpropagated content in the first translated content. Of course, the progress visualization can also be represented by other patterns or shapes, and this application does not impose any specific limitations on this.

[0114] For example, please refer to Figure 6 , Figure 6 This application provides a schematic diagram illustrating how a progress bar represents the interpretation progress. For example... Figure 6As shown, in the interpretation card, the visual identifier can be a progress bar. The colored part of the progress bar represents the duration of propagation (i.e., the playback duration of the propagated content), while the white part represents the duration of non-playback (i.e., the expected propagation duration of the non-propagated content). Therefore, this progress bar can be used to indicate the expected propagation duration of the non-propagated content in the first translation.

[0115] Of course, the progress visualization can also be a text indicator. That is, the first electronic device can also display the expected duration of the un-disseminated content in the first translated content through text. For example, the first electronic device could display the text "5 seconds remaining before playback," thereby indicating that the expected duration of the un-disseminated content is 5 seconds.

[0116] Furthermore, in some implementations, the first electronic device may simultaneously employ any two or three of the methods described in methods 1-3 to display the first translation progress. That is, the first electronic device may simultaneously display any two or three of the following three parts: The first part comprises the disseminated and / or undisseminated content of the first translation; the second part comprises the sentence disseminated and / or word disseminated progress; and the third part comprises the expected disseminated duration of the undisseminated content in the first translation.

[0117] For example, in Figure 5 In the interpretation card shown, the first electronic device simultaneously displays the transmitted content, the untransmitted content, and the number of untransmitted sentences, that is, it simultaneously displays the first part of the content and the second part of the content mentioned above.

[0118] Optionally, to further assist the speaker in controlling their speaking pace, when the untranslated content in the first translated content meets preset conditions, the first electronic device displays a prompt message to remind the speaker to slow down their speaking speed.

[0119] The aforementioned preset conditions include that the number of sentences included in the unspread content is greater than or equal to the target threshold, or that the spread duration of the unspread content is greater than or equal to the target duration. The target threshold can be set or adjusted according to the actual scenario, for example, a value of 2, 3, 4, or 5; the target duration can also be set or adjusted according to the actual scenario, for example, 30 seconds, 60 seconds, or 90 seconds. This application does not specifically limit the target threshold and target duration.

[0120] For example, in a specific scenario, when the number of sentences included in the untransmitted content is greater than or equal to two, or the duration of the untransmitted content's transmission is greater than or equal to 30 seconds, the first electronic device can display a prompt message to remind the speaker to slow down their speaking speed. For example, please refer to... Figure 7 , Figure 7This application provides a schematic diagram illustrating the display of prompt information. For example... Figure 7 As shown, in the interpretation card displayed on the meeting interface, the interpretation progress shows that two sentences remain unplayed, which meets the condition that the number of sentences included in the unplayed content is greater than or equal to two. Therefore, a prompt message is displayed in the upper right corner of the interpretation card, which reads: "Please slow down your speaking pace appropriately."

[0121] Furthermore, in some designs, the prompts can have multiple different levels, and the display method for each level is different. For example, different levels of prompts can be displayed using different colors or color depths. Alternatively, one level of prompt may be constantly displayed, while another level may be flashing. Specifically, the slower the first interpretation progress, the higher the level of the prompt can be, thus more prominently reminding the speaker to slow down.

[0122] The above describes how the first electronic device displays the first interpretation progress for the first translated content. However, in some conference scenarios, because participants come from different regions, the raw audio data captured by the first electronic device may be translated into content corresponding to different languages. In this case, since the length of the translated content may differ between languages, the interpretation progress for the translated content in different languages ​​may also be different. Based on this, this application proposes displaying the interpretation progress corresponding to different translated content on the first electronic device used by the speaker, so that the speaker can clearly know the interpretation progress for the translated content in various languages.

[0123] For example, the first electronic device can also acquire a second translation progress, which indicates the progress of transmitting the second translated content. The second translated content is the content obtained by translating the original speech data, and the language of the second translated content is different from the language of the first translated content. For example, the language of the original speech data is Chinese, the language of the first translated content is English, and the language of the second translated content is Japanese.

[0124] In this way, after obtaining the second interpretation progress, the first electronic device can display the second interpretation progress.

[0125] There are several ways in which the first electronic device can display the second interpretation progress.

[0126] In one possible implementation, the first electronic device simultaneously displays the first interpretation progress and the second interpretation progress. For example, if both the first and second interpretation progress are displayed as progress bars, the first electronic device may display multiple progress bars simultaneously, with different progress bars corresponding to different interpretation progresses (i.e., the aforementioned first and second interpretation progresses).

[0127] In this way, by displaying the interpretation progress for different languages ​​simultaneously, speakers can clearly see the interpretation progress for each language, thereby better controlling their speaking pace and ensuring the effectiveness of the meeting.

[0128] In another possible implementation, in response to receiving a display switching command, the first electronic device switches from displaying a first interpretation progress to displaying a second interpretation progress. That is, the first electronic device displays only the interpretation progress for one language; and, upon receiving a display switching command, the first electronic device can then switch to displaying the interpretation progress for other languages.

[0129] For example, please refer to Figure 8 , Figure 8 This is a schematic diagram illustrating a switching display of interpretation progress provided in this application. Figure 8 As shown, in Figure 8 On the left side of the conference interface, the interpretation cards display the interpretation progress corresponding to the English version. When a speaker clicks the corresponding Japanese icon using a mouse or touch, they can send a display switching command to the first electronic device. In response to this command, the first electronic device switches from displaying the English interpretation progress to displaying the Japanese interpretation progress.

[0130] Furthermore, when both the first and second interpretation progress can be indicated in multiple ways, the first electronic device can also simultaneously display the first and second interpretation progress in a simple display mode, and display one of the first and second interpretation progresses in a more detailed display mode.

[0131] For example, in Figure 7 In this context, the first electronic device can simultaneously display the translation progress for English, Japanese, and Korean using a bar chart (i.e., the simplified display method described above); furthermore, the first electronic device can also display the translation progress for English by showing the content that has been transmitted and the content that has not been transmitted (i.e., the detailed display method described above).

[0132] It should be noted that if multiple attendees in a meeting are using electronic devices that are all configured for the same translation language, but each device has a different playback speed, the interpretation progress of the same text may differ across these devices. In such cases, to ensure the simplicity of the display interface, the meeting interface can only show the interpretation progress for a specific attendee (e.g., the attendee with the slowest playback speed) within the same translation language. Of course, the meeting interface can also display the interpretation progress for all attendees; this application does not impose a specific limitation on this.

[0133] Optionally, in some scenarios, the first electronic device can also acquire and display the voiceprint of the original speech data. Here, a voiceprint is a sound wave spectrum for speech data. When the voiceprint of the original speech data is represented by a voiceprint graph, the first electronic device can convert the changes in sound waves in the original speech data into changes in the intensity, wavelength, frequency, and rhythm of electrical signals, and plot these changes in electrical signals as a spectral graph, thereby obtaining the voiceprint graph.

[0134] For example, please refer to Figure 9 , Figure 9 This is a schematic diagram illustrating a voiceprint display provided in this application. Figure 9 As shown, on the interpretation card, the first electronic device can display the voiceprint of the raw speech data captured by the first electronic device, and display it in the form of a spectral graph.

[0135] In summary, by displaying the voiceprint of the original speech data, this solution makes it easier for speakers to understand the rhythm of their previous speech (such as changes in volume, speed, and pauses), thus enabling them to adaptively adjust their subsequent speaking rhythm.

[0136] The above describes how the first electronic device displays the first interpretation progress. The following will describe how the first electronic device obtains the first interpretation progress.

[0137] Optionally, after the first electronic device captures the raw voice data, the first electronic device sends the raw voice data to a server, which performs translation on the raw voice data to obtain the first translated content.

[0138] Then, the first electronic device receives the first translated content and / or the expected propagation duration of the first translated content sent by the server. That is, after the server performs translation on the original voice data to obtain the first translated content, the server can send the first translated content, or the expected propagation duration of the first translated content, or even send both the first translated content and the expected propagation duration of the first translated content to the first electronic device.

[0139] In this way, the first electronic device can determine the first translation progress based on the first translated content and / or the expected transmission duration of the first translated content.

[0140] Understandably, when determining the first translation progress, the first electronic device can determine that at the time point when the server sends the first translated content and / or the expected propagation time of the first translated content to the first electronic device, the first translation progress is specifically no progress (i.e., the first translated content has just begun to be translated). Furthermore, as time changes, the specific progress of the first translation will gradually change.

[0141] In addition, after obtaining the first translation, the server can send the first translation to other electronic devices (such as the electronic devices used by the audience), so that the first translation can be played as audio on other electronic devices to complete the simultaneous interpretation process.

[0142] In the above embodiments, the first electronic device sends the raw audio data to the server, which then performs the translation process. In some possible embodiments, if the first electronic device has sufficient computing resources, it can also translate the raw audio data and send the resulting first translation to the server, which then forwards it to other electronic devices. Of course, if multiple listeners speaking different languages ​​are attending the meeting, the first electronic device can also translate the raw audio data into translations corresponding to different languages ​​and send the resulting multiple translations to the server, which then sends each translation to the corresponding listener's electronic device.

[0143] Optionally, after obtaining the first translated content, the server can send the first translated content and the digital human video to the second electronic device. The digital human video is used to demonstrate the digital human's speech process, and the changes in the digital human's mouth shape during speech are determined based on the first translated content.

[0144] In this way, after receiving the first translated content and the digital human video, the second electronic device can play the digital human video and the first translated content simultaneously, thereby achieving audio-visual synchronization.

[0145] In this context, a digital human is a digitized human figure created using digital technology that closely resembles a human. In this application, the digital human can be generated based on the speaker's image. Furthermore, the digital human video is a video used to demonstrate the speaker's delivery of a speech in a specific language.

[0146] Specifically, since the speaker speaks in a particular language, the original audio data is obtained. However, during simultaneous interpretation, the speaker's original audio data is translated into audio data in another language. Therefore, if the actual video of the speaker's speech is used as the display, there may be a mismatch between the lip movements and the translated audio data. Based on this, in this application, the server determines the lip movements of the digital human during speech based on the first translated content, so that the lip movements of the digital human match the speaking style of the first translated content itself.

[0147] In practical applications, if the server needs to translate the original voice data into multiple different languages, then the server also needs to generate corresponding digital human videos for each language.

[0148] Specifically, after translating the raw audio data to obtain the translated content corresponding to different languages, the server can first obtain a digital human model. This digital human model can be pre-created by the speaker on their electronic device, or it can be created in real-time by the server based on the speaker's video. That is, the conference system supports custom digital humans; speakers can pre-create their digital humans before joining the conference and automatically upload them to the server upon joining. Alternatively, speakers can record a video of themselves speaking and upload it to the server after joining the conference, where the server will generate a digital human model in real-time based on that video.

[0149] Then, the server remodels the digital human model based on the translated content corresponding to different languages, thereby obtaining digital human videos corresponding to different languages. This remodeling can actually use translated text or speech data to drive lip movements, remodeling the digital human's mouth shape data to obtain multiple digital humans arranged in chronological order. Each digital human, as a video frame, is then matched temporally with the translated speech data to obtain the digital human video. In other words, the server can have a built-in pre-set algorithm that, after generating translated text or speech data, uses this algorithm to drive lip movements in the digital human, resulting in multiple digital humans with different lip shapes. These digital humans with different lip shapes, then packaged into a video in chronological order, yield the digital human video.

[0150] Finally, the server sends the translated content and digital human video in the same language to the electronic devices used by the corresponding listeners. In this way, for listeners speaking different languages, the digital human they see has different lip movements, matching the lip movements required for speaking in their respective languages.

[0151] To facilitate understanding, the following will provide a detailed explanation of the specific process of implementing simultaneous interpretation in a conference, using concrete examples.

[0152] Please see Figure 10 , Figure 10 This application provides a flowchart illustrating a simultaneous interpretation process in a conference. For example... Figure 10 As shown, the conference system includes the electronic devices used by the speakers (hereinafter referred to as speaker electronic devices), a server, and electronic devices used by one or more audience members (hereinafter referred to as audience electronic devices). The process of implementing simultaneous interpretation in a conference includes the following steps 1001-1006.

[0153] Step 1001: The listener's electronic device uploads the language setting parameters to the server.

[0154] After joining the conference, attendees' electronic devices can upload language setting parameters to the server. This allows the server to subsequently send corresponding audio data to the attendees' electronic devices based on these parameters. Specifically, the conference system client provides a language setting option, which attendees can use on their electronic devices to generate language setting parameters. These parameters include a language type parameter, indicating the language to which the original audio data should be translated; that is, the language type of the translated audio data received by the attendee's electronic device. Additionally, language setting parameters may include a speech rate parameter, indicating the speed at which the translated audio data is played after being received by the attendee's electronic device. Furthermore, language setting parameters may also include tone parameters, pitch parameters, and other information.

[0155] It should be noted that when the conference system includes multiple audience electronic devices, different audience electronic devices may upload different language settings parameters to the server. For example, audience electronic device 1 may be set to English, audience electronic device 2 to Japanese, and audience electronic device 3 to Korean.

[0156] In this scenario, the listener's electronic device can proactively upload language settings parameters to the server after joining the meeting; alternatively, the listener's electronic device can include the language settings parameters in the connection request when establishing a connection with the server. Alternatively, the server can send a data request to the listener's electronic device after it joins the meeting, allowing the listener's electronic device to then send the language settings parameters back to the server based on the data request.

[0157] Optionally, in some scenarios, the speaker's electronic device can also upload language setting parameters to the server. Since audience members may also speak during the meeting (e.g., ask questions), their statements may need to be translated before being sent to the speaker's electronic device. Therefore, the speaker's electronic device can upload language setting parameters to the server, allowing the server to determine the language type of the audio data received from the audience's electronic device before returning it to the speaker's electronic device.

[0158] Step 1002: The speaker's electronic device sends the raw voice data to the server.

[0159] During the speaker's speech, the speaker's electronic devices can acquire the raw audio data of the speaker's speech through devices such as microphones and send the raw voice data to the server.

[0160] Step 1003: The server translates the original speech data to obtain translated speech data.

[0161] In the process of translating raw audio data, the server can first perform audio-to-text processing, that is, extract the corresponding text from the raw audio data according to a preset audio-to-text algorithm. Then, the server translates the extracted text according to the obtained language setting parameters, thus obtaining the translated text. Of course, if different listeners' electronic devices are set to different translation languages, then the server needs to translate the text in different languages. For the translated text, the server can use text-to-audio technology to convert the translated text into translated audio data.

[0162] Optionally, the conference system can include at least one virtual human. By using translated text to drive the virtual human, translated speech data corresponding to the text can be obtained. Different virtual humans can have different genders, voices, or voiceprints. Therefore, in step 1001 above, the listener's electronic device can also be configured with information such as the gender, voice, or voiceprint of the virtual human receiving the translated speech data.

[0163] In addition, the server can directly perform language conversion on the original speech data to obtain translated speech data, without having to translate the original speech data into translated text before performing the conversion.

[0164] Optionally, the server can first perform a series of preprocessing steps on the raw audio data, such as noise reduction and voice amplification, to obtain clearer audio data before starting the translation process described above.

[0165] It should be noted that when different listeners' electronic devices are set to different translation languages, the server can perform multiple translations on the original language data to obtain multiple translated audio data in different languages.

[0166] It is worth noting that when the server converts the raw audio data into translated audio data, it can perform the conversion separately for each listener, meaning the number of translated audio data points is consistent with the number of listeners. Alternatively, the server can determine the number of audio data points to be converted based on the obtained language setting parameters.

[0167] For example, in a conference system with 5 participants (1 speaker and 4 listeners), the speaker speaks in Chinese, and the server obtains the language settings for the listeners as Japanese, English, English, and Korean. During the translation phase, the server can perform only 3 text translations, resulting in translated texts in Japanese, English, and Korean. Then, when the server uses the translated text to drive a virtual human to obtain translated speech data, it can correspond to the 4 listeners, thus obtaining 4 sets of translated speech data.

[0168] In another scenario, depending on the language settings of the listeners, if none of the listeners have personalized their virtual persona or speech rate (or the conference system does not support personalized settings), the server can use the default virtual persona to convert the three translated texts into audio, resulting in three translated audio data sets. The translated audio data corresponding to the English text needs to be sent to two different listener electronic devices.

[0169] For example, please refer to Figure 11 , Figure 11 This is a schematic diagram illustrating the conversion of translated text into translated speech data, as provided in this application. Figure 11 As shown, in Figure 11 In the conversion process shown on the left, the Chinese text is translated into three texts: English, Japanese, and Korean. Each text is then converted into corresponding audio data. For example, the English text is converted into English audio, the Japanese text into Japanese audio, and the Korean text into Korean audio. Furthermore, the English audio needs to be sent to listener 1 and listener 2, the Japanese audio to listener 2, and the Korean audio to listener 3. Here, "listener" refers to the meeting client running on the listener's electronic device.

[0170] exist Figure 11In the conversion process shown on the right, the Chinese text is translated into three texts: English, Japanese, and Korean. The English text is then converted into two corresponding audio data sets. For example, the English text is converted into English Audio 1 and English Audio 2, the Japanese text into Japanese audio, and the Korean text into Korean audio. Furthermore, English Audio 1 needs to be sent to listener 1, English Audio 2 to listener 2, the Japanese audio to listener 3, and the Korean audio to listener 4.

[0171] Step 1004: The server generates a digital human video based on the translated speech data.

[0172] After obtaining the translated speech data that needs to be distributed to the electronic devices of various listeners, the server can further generate digital human videos based on the translated speech data, thereby obtaining digital human videos that match the lip movements with the translated speech data.

[0173] Step 1005: The server sends the translated voice data and digital human video to the listener's electronic device, and sends the translated text and / or translation progress corresponding to the translated voice data to the speaker's electronic device.

[0174] Specifically, in order for the speaker's electronic device to display the translation progress, the server can send the translated text corresponding to the translated audio data and / or translation progress information to the speaker's electronic device. The translation progress information sent by the server can include, for example, the expected playback duration of the translated audio data, the currently transmitted sentence, etc., allowing the speaker's electronic device to determine the translation progress based on this information.

[0175] Step 1006: The speaker's electronic device displays the interpretation progress, and the audience's electronic devices play the translated audio data and digital human video.

[0176] Specifically, after receiving the interpretation progress information, the speaker's electronic device can determine the interpretation progress corresponding to each audience member's electronic device and display the corresponding interpretation progress on the conference interface. Simultaneously, the audience's electronic devices can play translated audio data and digital human videos to complete the simultaneous interpretation process.

[0177] The methods provided in the embodiments of this application have been described in detail above. Next, the device for performing the above methods provided in the embodiments of this application will be described.

[0178] Please see Figure 12 , Figure 12 This is a schematic diagram of an interpretation progress display device provided in this application. Figure 12As shown, the interpretation progress display device is deployed on the first electronic device. The interpretation progress display device includes: an acquisition module 1201, used to acquire the first interpretation progress, which is used to indicate the propagation progress of the first translated content, which is the content obtained by translating the original voice data captured by the first electronic device; and a display module 1202, used to display the first interpretation progress.

[0179] In one possible implementation, the display module 1202 is specifically used to display the propagated content and / or unpropagated content in the first translated content.

[0180] In one possible implementation, the display module 1202 is specifically used to: display a progress visualization indicator, which is used to indicate the progress of the first interpretation.

[0181] In one possible implementation, progress visualization markers are used to indicate the progress of sentence propagation and / or word propagation of the first translated content.

[0182] In one possible implementation, progress visualization markers are used to indicate the expected propagation time of unpropagated content in the first translation.

[0183] In one possible implementation, the display module 1202 is further configured to: display a prompt message when the untranslated content in the first translated content meets a preset condition, the prompt message being used to prompt the speaker to reduce the speaking speed.

[0184] In one possible implementation, the preset conditions include that the number of statements included in the unspread content is greater than or equal to the target threshold, or that the spread duration of the unspread content is greater than or equal to the target duration.

[0185] In one possible implementation, the first translation content is translated text, and the language of the translated text is different from the language of the original speech data; or, the first translation content is translated speech data, and the language of the translated speech data is different from the language of the original speech data.

[0186] In one possible implementation, the acquisition module 1201 is further configured to acquire a second translation progress, which indicates the progress of transmitting the second translated content. The second translated content is the content obtained by translating the original speech data, and the language corresponding to the second translated content is different from the language corresponding to the first translated content. The display module 1202 is further configured to display the second translation progress.

[0187] In one possible implementation, the display module 1202 is specifically used to: simultaneously display the first interpretation progress and the second interpretation progress; or, in response to receiving a display switching instruction, switch from displaying the first interpretation progress to displaying the second interpretation progress.

[0188] In one possible implementation, the acquisition module 1201 is further used to acquire the voiceprint of the original voice data; the display module 1202 is further used to display the voiceprint.

[0189] In one possible implementation, a first electronic device is located in a conference system, which translates the raw voice data into first translated content and transmits the first translated content to a second electronic device in the conference system, which plays the first translated content via voice.

[0190] In one possible implementation, the acquisition module 1201 is specifically used for: sending raw voice data to the server; receiving the first translated content and / or the expected propagation duration of the first translated content sent by the server; and determining the first translation progress based on the first translated content and / or the expected propagation duration of the first translated content.

[0191] Please see Figure 13 , Figure 13 A schematic diagram of another interpretation progress display device provided in this application. Figure 13 As shown, the interpretation progress display device is deployed on a server. The interpretation progress display device includes: a transceiver module 1301, used to receive raw voice data sent by a first electronic device; a processing module 1302, used to perform translation on the raw voice data to obtain first translated content; the transceiver module 1301 is also used to send the first translated content and / or the expected propagation time of the first translated content to the first electronic device, the first translated content and / or the expected propagation time of the first translated content being used to enable the first electronic device to display the first interpretation progress, the first interpretation progress being used to indicate the progress of propagating the first translated content.

[0192] In one possible implementation, the transceiver module 1301 is further configured to: send the first translated content and the digital human video to the second electronic device; wherein the digital human video is used to display the digital human's speaking process, and the changes in the digital human's mouth shape during the speaking process are determined based on the first translated content.

[0193] Please see Figure 14 , Figure 14 This is a schematic diagram of an execution device provided in this application. The execution device 1400 can specifically be a mobile phone, tablet, laptop, smart wearable device, server, etc., and is not limited thereto. Specifically, the execution device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403, and a memory 1404 (wherein the execution device 1400 may have one or more processors 1403). Figure 14(Taking a processor as an example), processor 1403 may include application processor 14031 and communication processor 14032. In some embodiments of this application, receiver 1401, transmitter 1402, processor 1403 and memory 1404 may be connected via a bus or other means.

[0194] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0195] Processor 1403 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses are referred to as the bus system in the diagram.

[0196] The methods disclosed in the embodiments of this application described above can be applied to, or implemented by, processor 1403. Processor 1403 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by integrated logic circuits in the hardware of processor 1403 or by instructions in software form. Processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0197] The processor 1403 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 1404. The processor 1403 reads information from memory 1404 and, in conjunction with its hardware, completes the steps of the above methods.

[0198] Receiver 1401 can be used to receive input digital or character information, and to generate signal inputs related to the settings and function control of the execution device. Transmitter 1402 can be used to output digital or character information through the first interface; transmitter 1402 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; transmitter 1402 may also include a display device such as a display screen.

[0199] This application also provides a chip, comprising a processing unit and a communication unit. The processing unit may be, for example, a processor, and the communication unit may be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in a storage unit to cause the chip within the execution device to perform the methods described in the above embodiments. Optionally, the storage unit can be an in-chip storage unit, such as a register or cache. Alternatively, the storage unit can be an external storage unit located within a wireless access device, such as a read-only memory (ROM) or other types of static storage devices capable of storing static information and instructions, such as random access memory (RAM).

[0200] Please refer to Figure 15 , Figure 15 This is a schematic diagram of a computer-readable storage medium provided in this application. This application also provides a computer-readable storage medium in some embodiments, wherein the above-described... Figure 3 The disclosed method can be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or on other non-transitory media or articles of art.

[0201] Figure 15 A conceptual partial view of an example computer-readable storage medium arranged according to at least some of the embodiments shown herein is illustrated schematically. The example computer-readable storage medium includes a computer program for executing computer processes on a computing device.

[0202] In one embodiment, the computer-readable storage medium 1500 is provided using a signal bearer medium 1501. The signal bearer medium 1501 may include one or more program instructions 1502, which, when executed by one or more processors, can provide the above-mentioned... Figure 3 The described function or part of the function.

[0203] In some examples, the signal carrying medium 1501 may include a computer-readable medium 1503, such as, but not limited to, a hard disk drive, a compact disc (CD), a digital video disc (DVD), a digital magnetic tape, a memory, ROM, or RAM, etc.

[0204] In some embodiments, the signal-bearing medium 1501 may include a computer-recordable medium 1504, such as, but not limited to, a memory, a read / write (R / W) CD, a R / W DVD, etc. In some embodiments, the signal-bearing medium 1501 may include a communication medium 1505, such as, but not limited to, digital and / or analog communication media (e.g., fiber optic cables, waveguides, wired communication links, wireless communication links, etc.). Therefore, for example, the signal-bearing medium 1501 may be transmitted by a wireless communication medium 1505 (e.g., a wireless communication medium conforming to the IEEE 802.X standard or other transmission protocols).

[0205] One or more program instructions 1502 may be, for example, computer-executable instructions or logical implementation instructions. In some examples, the computing device may be configured to provide various operations, functions, or actions in response to one or more program instructions 1502 conveyed to the computing device via a computer-readable medium 1503, a computer-recordable medium 1504, and / or a communication medium 1505.

[0206] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0207] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods of the various embodiments of this application.

[0208] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product.

[0209] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, computer instructions may be transferred from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

Claims

1. A method for displaying interpretation progress, characterized in that, include: The first electronic device acquires a first translation progress, which is used to indicate the propagation progress of the first translated content, which is the content obtained by translating the original speech data captured by the first electronic device. The first electronic device displays the first translation progress.

2. The method according to claim 1, characterized in that, The first electronic device displays the first translation progress, including: The first electronic device displays the disseminated and / or undisseminated content from the first translated content.

3. The method according to claim 1, characterized in that, The first electronic device displays the first translation progress, including: The first electronic device displays a progress visualization indicator, which is used to indicate the progress of the first interpretation.

4. The method according to claim 3, characterized in that, The progress visualization indicator is used to indicate the progress of sentence propagation and / or word propagation of the first translated content.

5. The method according to claim 3 or 4, characterized in that, The progress visualization indicator is used to indicate the expected propagation time of the unpropagated content in the first translated content.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: When the untranslated content in the first translated content meets the preset conditions, the first electronic device displays a prompt message, which is used to prompt the speaker to slow down their speaking speed.

7. The method according to claim 6, characterized in that, The preset conditions include the number of statements included in the unspread content being greater than or equal to the target threshold, or the spread duration of the unspread content being greater than or equal to the target duration.

8. The method according to any one of claims 1-7, characterized in that, The first translation content is translated text, and the language corresponding to the translated text is different from the language corresponding to the original speech data; Alternatively, the first translation content may be translated speech data, and the language corresponding to the translated speech data may be different from the language corresponding to the original speech data.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: The first electronic device acquires a second translation progress, which is used to indicate the progress of transmitting the second translated content. The second translated content is the content obtained by translating the original speech data, and the language corresponding to the second translated content is different from the language corresponding to the first translated content. The first electronic device displays the second translation progress.

10. The method according to claim 9, characterized in that, The first electronic device displays the second translation progress, including: The first electronic device simultaneously displays the first translation progress and the second translation progress; Alternatively, in response to receiving a display switching instruction, the first electronic device switches from displaying the first translation progress to displaying the second translation progress.

11. The method according to any one of claims 1-10, characterized in that, The method further includes: The first electronic device acquires the voiceprint of the original voice data; The first electronic device displays the voiceprint.

12. The method according to any one of claims 1-11, characterized in that, The first electronic device is located in a conference system. The conference system is used to translate the original voice data into the first translated content and transmit the first translated content to a second electronic device in the conference system. The second electronic device is used to play the first translated content via voice.

13. The method according to any one of claims 1-12, characterized in that, The first electronic device acquires the first translation progress, including: The first electronic device sends the raw voice data to the server; The first electronic device receives the first translated content sent by the server and / or the expected propagation duration of the first translated content; The first electronic device determines the first translation progress based on the first translated content and / or the expected transmission duration of the first translated content.

14. A method for displaying interpretation progress, characterized in that, include: The server receives the raw voice data sent by the first electronic device; The server translates the original speech data to obtain the first translated content; The server sends the first translated content and / or the expected propagation duration of the first translated content to the first electronic device. The first translated content and / or the expected propagation duration of the first translated content are used to cause the first electronic device to display a first translation progress, which is used to indicate the progress of propagating the first translated content.

15. The method according to claim 14, characterized in that, The method further includes: The server sends the first translated content and the digital human video to the second electronic device; The digital human video is used to demonstrate the digital human's speaking process, and the changes in the digital human's mouth shape during the speaking process are determined based on the first translated content.

16. A translation progress display device, characterized in that, The device is deployed on a first electronic device, the device comprising: The acquisition module is used to acquire the first translation progress, which is used to indicate the propagation progress of the first translated content, which is the content obtained by translating the original voice data captured by the first electronic device. The display module is used to display the progress of the first interpretation.

17. The apparatus according to claim 16, characterized in that, The display module is specifically used for: Displays the disseminated and / or undisseminated content from the first translated content.

18. The apparatus according to claim 16, characterized in that, The display module is specifically used for: Display a progress visualization indicator, which is used to indicate the progress of the first interpretation.

19. The apparatus according to claim 18, characterized in that, The progress visualization indicator is used to indicate the progress of sentence propagation and / or word propagation of the first translated content.

20. The apparatus according to claim 18 or 19, characterized in that, The progress visualization indicator is used to indicate the expected propagation time of the unpropagated content in the first translated content.

21. The apparatus according to any one of claims 16-20, characterized in that, The display module is also used for: When the untranslated content in the first translated content meets the preset conditions, a prompt message is displayed, which is used to prompt the speaker to slow down their speaking speed.

22. The apparatus according to claim 21, characterized in that, The preset conditions include the number of statements included in the unspread content being greater than or equal to the target threshold, or the spread duration of the unspread content being greater than or equal to the target duration.

23. A translation progress display device, characterized in that, The device is deployed on a server, and the device includes: A transceiver module is used to receive raw voice data sent by the first electronic device; The processing module is used to translate the original speech data to obtain the first translated content; The transceiver module is further configured to send the first translated content and / or the expected propagation duration of the first translated content to the first electronic device, wherein the first translated content and / or the expected propagation duration of the first translated content are used to enable the first electronic device to display a first translation progress, wherein the first translation progress is used to indicate the progress of propagating the first translated content.

24. The apparatus according to claim 23, characterized in that, The transceiver module is also used for: Send the first translated content and the digital human video to the second electronic device; The digital human video is used to demonstrate the digital human's speaking process, and the changes in the digital human's mouth shape during the speaking process are determined based on the first translated content.

25. An electronic device, characterized in that, The device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code, wherein when the code is executed, the device performs the method as described in any one of claims 1 to 13.

26. A server, characterized in that, The device includes a memory and a processor; the memory stores code, and the processor is configured to execute the code, wherein when the code is executed, the device performs the method as described in any one of claims 14 to 15.

27. A conference system, characterized in that, This includes the electronic device as described in claim 25 and the server as described in claim 26.

28. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed by the computer, cause the computer to perform the method according to any one of claims 1 to 15.

29. A computer program product, characterized in that, The computer program product stores instructions that, when executed by a computer, cause the computer to perform the method described in any one of claims 1 to 15.