Communication method, communication system and electronic device

By pre-storing and sending the translated audio data at the sending end, the latency problem in simultaneous interpretation solutions is solved, enabling faster information reception and more efficient online communication.

WO2026032369A1PCT designated stage Publication Date: 2026-02-12HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/113210
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-08-07
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

In scenarios such as international conferences, remote work, and instant messaging where rapid response and real-time communication are required, simultaneous interpretation solutions suffer from latency issues and cannot meet users' needs for fast, real-time language conversion services.

Method used

The electronic device at the transmitting end receives user input in advance and stores the translated voice data. After detecting the user's sending operation, it sends the translated voice data to the receiving end, reducing data transmission latency.

Benefits of technology

By pre-storing and sending translated voice data, latency during communication is reduced, providing a smooth and real-time online communication experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025113210_12022026_PF_FP_ABST
    Figure CN2025113210_12022026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the embodiments of the present application are a communication method, a communication system and an electronic device. A transmitter electronic device can receive a user input in advance, and, on the basis of the user input, generate a corresponding translation voice and store same; upon detecting a sending operation of the user, the transmitter electronic device can send the pre-stored translation voice to a receiver electronic device for playback, such that the delay of data transmission during communication is greatly reduced, thereby reducing the online waiting time for users, and further providing smooth and real-time online communication experience for users.
Need to check novelty before this filing date? Find Prior Art

Description

Communication method, communication system and electronic device

[0001] The present application claims priority to the Chinese patent application No. 202411096168.7, filed on August 9, 2024, with the State Intellectual Property Office of China, and entitled "A communication method, a communication system and an electronic device", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] The present application relates to the technical field of terminal, and in particular to a communication method, a communication system and an electronic device. BACKGROUND

[0003] In some call or conference scenarios, users often need to communicate and exchange information online with other users using different languages, which poses a certain challenge and trouble to users who cannot use foreign languages fluently.

[0004] Although the simultaneous interpretation solution can alleviate the problem of cross-language communication to a certain extent, it still has the problem of delay in scenarios such as international conferences, remote work and instant messaging that require quick response and real-time communication, and cannot meet the user's demand for quick and real-time language conversion services. SUMMARY

[0005] The present application provides a communication method, a communication system and an electronic device. The electronic device of the sending end can pre-receive user input, generate and store the corresponding translation voice based on the user's input. After detecting the sending operation of the user, the electronic device of the sending end can send the pre-stored translation voice to the receiving end electronic device for playing, thereby greatly reducing the time delay of data transmission in the communication process, i.e. reducing the user's online waiting time, and thus providing a smooth and real-time online communication experience for the user.

[0006] In a first aspect, the present application provides a communication method applied to a first electronic device, the method comprising: starting a first communication application program by the first electronic device, establishing a first session in the first communication application program by a first user using the first electronic device and a second user using a second electronic device; detecting a first operation by the first electronic device, displaying a first interface, and displaying one or more texts in the first interface, and storing voice data of a translation of the text by the first electronic device; detecting a user operation of sending the text during the existence of the first session, and sending the voice data of the translation of the text to the second electronic device by the first electronic device.

[0007] The method provided by the first aspect is implemented, and in the online communication scenario of the first communication application of the first electronic device (i.e., the electronic device of the sending end) and the second electronic device (i.e., the electronic device of the receiving end), the first electronic device (i.e., the electronic device of the sending end) can enter the first interface by clicking the floating ball or the first operation, and the first interface already contains one or more texts. During the online communication, the electronic device of the sending end can first store the voice data of the translation of the text, and send the voice data of the translation of the text to the second electronic device after detecting that the user performs the sending operation.

[0008] Compared with the simultaneous interpretation scheme, the scheme of first storing the voice data of the translation of the text in the electronic device of the sending end and sending the voice data of the translation of the text to the electronic device of the receiving end after detecting the sending operation of the text can obtain the voice data of the translation of the text in advance, so that the translation processing does not need to be performed in real time when the user speaks, which greatly reduces the time delay of the voice data of the translation of the text from the sending end to the receiving end, and also means that the user of the receiving end no longer needs to wait for a long time for the electronic device of the sending end to complete the voice processing and sending, thereby realizing faster information receiving and more efficient online communication.

[0009] In combination with the method provided by the first aspect, in some embodiments, the first interface displays one or more texts, and before the first electronic device stores the voice data of the translation of the text, the method further includes: the first electronic device displays an input box in the first interface; the first electronic device receives the text input by the user from the input box; and the first electronic device obtains the voice data of the translation of the text.

[0010] In combination with the method provided by the first aspect, in some embodiments, the first interface displays one or more texts, and before the first electronic device stores the voice data of the translation of the text, the method further includes: the first electronic device displays an input box in the first interface; the first electronic device receives the voice input of the user from the input box; the first electronic device converts the voice input into text; and the first electronic device obtains the voice data of the translation of the text.

[0011] In the method provided by the above embodiments, the electronic device of the sending end can receive the voice input of the user through the input box, which is more in line with the habits of the user in online communication and is also convenient for the user to operate.

[0012] In combination with the method provided by the first aspect, in some embodiments, the text includes the original text and the translation of the text.

[0013] In the method provided by the above embodiments, the text displayed by the electronic device of the sending end in the first interface includes not only the original text but also the translated text. The user can see the original text and the translated text at the same time, which helps the user to better judge the accuracy of the translation.

[0014] In some embodiments of the method provided in the first aspect, the first electronic device stores speech data of the translation of the text, specifically including that the first electronic device stores spliced speech data of the text, the spliced speech data being spliced from speech data of the original text and speech data of the translation of the text, the speech data of the original text being obtained by a microphone of the first electronic device or being generated by the first electronic device based on the original text.

[0015] Implementing the method provided in the above embodiments, the electronic device of the sending end can not only send speech data of the translation, but also send speech data of the original text, and the sending is in the form of splicing the two pieces of speech data, so as to take care of both users using the original language and users using the translation in the online communication scenario. It can be understood that if the user inputs the text by voice, the speech data of the original text is the speech data picked up by the microphone of the electronic device of the sending end; if the user inputs the text by keyboard or imports the text from a memo or other sources, the speech data of the original text can be generated by the electronic device of the sending end using the text-to-speech technology.

[0016] In some embodiments of the method provided in the first aspect, before detecting the user operation of sending the text, the method further includes that after receiving the text input by the user, detecting a user operation of temporarily storing the text, and the first electronic device stores speech data of the translation of the text.

[0017] Implementing the method provided in the above embodiments, after receiving the text input by the user, if the electronic device of the sending end detects a user operation of temporarily storing the text, for example, the electronic device of the sending end detects that the user operates (for example, clicks) the temporary storage control, the electronic device of the sending end stores the speech data of the translation of the text in the local device first, so that when the user needs to send the speech data of the translation of the text, the electronic device of the sending end can directly obtain the speech data of the translation of the text from the device and send it to the electronic device of the receiving end, thereby reducing the delay of the speech data of the translation from the sending end to the receiving end, which means that the user of the receiving end no longer needs to wait for a long time for the electronic device of the sending end to complete the processing and sending of the speech, thereby realizing faster information reception and more efficient online communication.

[0018] In some embodiments of the method provided in the first aspect, after detecting the user operation of sending the text, the method further includes detecting a user operation of pausing sending the text, and the first electronic device stops sending the first part of the speech data to the second electronic device, the speech data of the translation of the text being composed of the first part of the speech data and a second part of the speech data, the second part of the speech data having been sent to the second electronic device by the first electronic device before the user operates the pause control.

[0019] According to the method provided in the above embodiment, when the electronic device of the sending end detects a user operation of pausing the sending of the text, the electronic device of the sending end can display a pause control. After the pause control is operated by the user, the electronic device of the sending end can stop sending the part (i.e., the first part of the voice data) of the voice data of the translation of the text being sent to the electronic device of the receiving end, thereby simulating the scenario of a sentence being interrupted in face-to-face conversation, and providing a more realistic and practical online communication experience for the user.

[0020] According to the method provided in the first aspect, in some embodiments, before detecting the user operation of sending the text, the method further includes: detecting a user operation of modifying the text, the first electronic device displaying the text in an input box in the first interface; the first electronic device receiving the text modified by the user from the input box; the first electronic device obtaining and storing the voice data of the translation of the modified text; detecting the user operation of sending the text, the first electronic device sending the voice data of the translation of the text to the second electronic device, specifically including: detecting the user operation of sending the modified text, the first electronic device sending the voice data of the translation of the modified text to the second electronic device.

[0021] According to the method provided in the above embodiment, the electronic device of the sending end can also modify the text. For example, the electronic device of the sending end can modify the text by operating the modification control by the user, thereby facilitating the user to modify the voice being sent and the voice to be sent, and thereby ensuring the accuracy of the voice sent by the user. Optionally, the first electronic device can also modify the voice that has been sent, thereby facilitating the user to correct some incorrect expressions in the previous communication.

[0022] According to the method provided in the first aspect, in some embodiments, after detecting the user operation of sending the text, the first electronic device sends the voice data of the translation of the text to the second electronic device, specifically including: the first electronic device detects a user operation of continuously sending a plurality of texts; and the first electronic device sends the voice data of the translation of the plurality of texts to the second electronic device in the order of the sending of the plurality of texts.

[0023] According to the method provided in the above embodiment, if the user continuously sends a plurality of texts, the electronic device of the sending end can send the voice data of the translation of the texts to the electronic device of the receiving end in the order of the sending of the plurality of texts, thereby reducing the time delay problem in the data transmission process.

[0024] According to the method provided in the first aspect, in some embodiments, after the first electronic device detects the user operation of continuously sending a plurality of texts, the method further includes: accompanying the continuous sending of the plurality of texts, the first electronic device labels the sending order of each of the plurality of texts.

[0025] Implementing the method provided by the above embodiment, the electronic device of the sending end can sequentially mark the plurality of texts sent continuously, thereby helping the user to determine whether the current text organization sequence is appropriate, and if the current text organization sequence is inappropriate, the user can cancel sending or pause sending.

[0026] In combination with the method provided by the first aspect, in some embodiments, the one or more texts displayed in the first interface include predefined texts.

[0027] Implementing the method provided by the above embodiment, the texts displayed in the first interface include predefined texts. The above predefined texts can be predefined by the electronic device of the sending end or predefined by the user. Thus, for some frequently used sentences, the user can not have to input through the input box every time.

[0028] In combination with the method provided by the first aspect, in some embodiments, the method further includes: after detecting a merging operation of the first text in the first interface and the second text in the first interface, the first electronic device splices the first text and the second text into a third text, speech data of a translation of the third text is obtained by splicing speech data of a translation of the first text and speech data of a translation of the second text, the plurality of texts include the first text and the second text, and the merging operation includes a user operation of dragging the first text to the second text.

[0029] Implementing the method provided by the above embodiment, the two texts in the first interface can be merged, and the corresponding speech data can also be merged, thereby simplifying the operation of the user through the merging and integrating manner.

[0030] In combination with the method provided by the first aspect, in some embodiments, the method further includes: the first electronic device receives speech data of the original text of the fourth text sent by the second electronic device, and plays the speech data of the original text of the fourth text sent by the second electronic device.

[0031] Implementing the method provided by the above embodiment, the electronic device of the sending end not only has the function of sending speech data, but also can be converted into the receiving end to receive and play speech data in some scenarios, such as the scenario of speech of other users in a meeting.

[0032] In combination with the method provided by the first aspect, in some embodiments, the method further includes: the first electronic device detects that the first electronic device has a dual-channel playback capability; the first electronic device translates the original text of the fourth text to obtain a translation of the fourth text; the first electronic device plays the speech data of the original text of the fourth text through one channel and plays the speech data of the translation of the fourth text through another channel.

[0033] The dual-channel playback capability can be provided by a single device (e.g., a Bluetooth headset) or by two devices (i.e., one device plays voice data of one channel).

[0034] The method provided in the above embodiments can be implemented. If the first electronic device has the dual-channel playback capability, the first electronic device can translate the original text of the fourth text sent by the second electronic device to obtain a translation of the fourth text, and generate voice data of the translation of the fourth text based on the translation, and then play the voice data of the original text in one channel of the first electronic device and play the voice data of the translation in another channel. This enables the first electronic device to play the voice data of the original text and the voice data of the translation at the same time, so that the user can choose to listen to the original text or the translation according to his language ability and preference, thereby providing a more personalized online communication experience for the user.

[0035] In combination with the method provided in the first aspect, in some embodiments, the method further includes: if the first electronic device receives the voice data of the translation of the fourth text sent by the second electronic device in the same tone after receiving the voice data of the original text of the fourth text sent by the second electronic device, the first electronic device determines not to play the voice data of the translation of the fourth text sent by the second electronic device.

[0036] The method provided in the above embodiments can be implemented. In the case where the second electronic device can send spliced voice data, if the first electronic device receives the voice data of the translation in the same tone, the first electronic device can determine not to play the repeated voice data again, thereby optimizing the experience of online communication of the user.

[0037] In the second aspect, the present application provides a communication method applied to a second electronic device, which includes: the second electronic device receives voice data of an original text of a text sent by a first electronic device, and plays the voice data of the original text of the text sent by the second electronic device; the second electronic device detects that it has a dual-channel playback capability; the second electronic device translates the original text to obtain a translation of the text; and the second electronic device plays the voice data of the original text through one channel and plays the voice data of the translation through another channel.

[0038] The method provided in the second aspect can be implemented. As a receiving-end electronic device, the second electronic device can translate the original text of the text sent by the first electronic device to obtain a translation of the text when it has a dual-channel playback capability, and generate voice data of the translation of the text based on the translation, and then play the voice data of the original text in one channel of the second electronic device and play the voice data of the translation in another channel. This enables the second electronic device to play the voice data of the original text and the voice data of the translation at the same time, so that the user can choose to listen to the original text or the translation according to his language ability and preference, thereby providing a more personalized online communication experience for the user.

[0039] In some embodiments, the method further comprises: if, after receiving the speech data of the original text sent by the first electronic device, the second electronic device continues to receive speech data of the translation of the text sent by the first electronic device in the same tone, the second electronic device determines not to play the speech data of the translation of the text sent by the first electronic device.

[0040] In the case that the first electronic device is capable of sending spliced speech data, if the second electronic device receives speech data of the translation in the same tone, the second electronic device can determine not to play the repeated speech data again, thereby optimizing the experience of online communication of the user.

[0041] In a third aspect, the present application provides an electronic device, which is the first electronic device, comprising a memory, a processor and a computer program stored in the memory; the processor executes the computer program to implement the method described in the first aspect and any possible implementation manner of the first aspect.

[0042] In a fourth aspect, the present application provides an electronic device, which is the first electronic device, comprising a memory, a processor and a computer program stored in the memory; the processor executes the computer program to implement the method described in the second aspect and any possible implementation manner of the second aspect.

[0043] In a fifth aspect, the present application provides a communication system comprising a first electronic device and a second electronic device, wherein the first electronic device is the electronic device described in the third aspect, and the second electronic device is the electronic device described in the fourth aspect.

[0044] In a sixth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method described in the first aspect and any possible implementation manner of the first aspect.

[0045] In a seventh aspect, the present application provides a computer program product comprising a computer program, and the computer program is executed by a processor to implement the method described in the first aspect and any possible implementation manner of the first aspect.

[0046] It can be understood that the electronic device provided in the third aspect, the electronic device provided in the fourth aspect, the computer storage medium provided in the sixth aspect and the computer program product provided in the seventh aspect are all used to execute the method provided in the present application. The communication system of the fifth aspect is composed of the electronic device provided in the third aspect and the electronic device provided in the fourth aspect. Therefore, the beneficial effects achieved by the above can refer to the beneficial effects in the corresponding method, which will not be described here. BRIEF DESCRIPTION OF DRAWINGS

[0047] Fig. 1 is a schematic diagram of an implementation process of simultaneous interpretation according to an embodiment of the present application;

[0048] Fig. 2 is a schematic diagram of a use interface of a call application according to an embodiment of the present application;

[0049] Fig. 3 is a schematic diagram of a use interface of an online conference application according to an embodiment of the present application;

[0050] Fig. 4 is a schematic diagram of a functional area distribution of a first interface according to an embodiment of the present application;

[0051] Fig. 5 is a schematic diagram of a functional area distribution of another first interface according to an embodiment of the present application;

[0052] Fig. 6 is a schematic diagram of a functional area distribution of another first interface according to an embodiment of the present application;

[0053] Fig. 7 is a schematic diagram of an uplink processing process in a source language speaking mode according to an embodiment of the present application;

[0054] Fig. 8 is a schematic diagram of an uplink processing process in a speaking and interpreting mode according to an embodiment of the present application;

[0055] Fig. 9 is a schematic diagram of an uplink processing process in a source language and target language speaking mode according to an embodiment of the present application;

[0056] Fig. 10 is a schematic diagram of an entry of a first interface according to an embodiment of the present application;

[0057] Fig. 11 is a schematic diagram of a first interface in a form of a floating window according to an embodiment of the present application;

[0058] Fig. 12 is a schematic diagram of another first interface in a form of a floating window according to an embodiment of the present application;

[0059] Fig. 13 is a schematic diagram of an interface before merging of a sentence control according to an embodiment of the present application;

[0060] Fig. 14 is a schematic diagram of an interface after merging of a sentence control according to an embodiment of the present application;

[0061] Fig. 15 is a schematic diagram of a downlink processing process according to an embodiment of the present application;

[0062] Fig. 16 is a schematic diagram of another downlink processing process according to an embodiment of the present application;

[0063] Fig. 17 is a schematic diagram of a structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0064] The terms used in the embodiments of the present application are for the purpose of describing particular embodiments only and are not intended to be limiting of the present application.

[0065] In some call or conference scenarios, users often need to communicate online with other users using different languages, which poses a certain challenge and trouble to users who cannot use foreign languages proficiently.

[0066] The electronic device can use simultaneous interpretation to solve the above problem. It can be understood that simultaneous interpretation is an efficient translation method. The electronic device can use speech recognition technology and translation algorithms to translate speech in different languages in real time, so that people with different language backgrounds can understand each other's meaning in real time, thereby overcoming language barriers and greatly improving the efficiency and quality of cross-language communication. FIG. 1 is a schematic diagram of an implementation process of simultaneous interpretation according to an embodiment of the present application.

[0067] Specifically, as shown in FIG. 1, the electronic device can pick up the user's voice through the microphone, and the language used by the user can be denoted as language 1. Then, the electronic device can use automatic speech recognition (ASR) technology to convert the user's voice into text, and translate the text recognized by ASR to obtain text in language 2. Further, the electronic device can use text-to-speech (TTS) technology to generate corresponding voice in language 2 based on the above text in language 2. In some embodiments, after generating the voice in language 2, the electronic device can send the voice in language 2 to the first communication application. Specifically, the electronic device provides the voice in language 2 to the software module responsible for the call function in the first communication application. In some embodiments, the first communication application has at least one of automatic speech recognition (ASR) and text-to-speech (TTS) functions. The above first communication application is an application that can perform online voice communication, including but not limited to a call application as shown in FIG. 2 and a conference application as shown in FIG. 3. After receiving the voice in language 2, the electronic device at the sending end can send the voice in language 2 to the first communication application on the electronic device at the receiving end through the first communication application, so that the first communication application on the electronic device at the receiving end can call the audio module to play the voice in language 2.

[0068] However, in the above-mentioned simultaneous interpretation process, the user cannot obtain a smooth online communication experience. One reason for the above-mentioned problem is that the sending end has a large processing delay for the picked-up voice. It can be understood that the processing delay of the sending end for the voice can include at least one of the following three parts: the delay of ASR, the delay of translation, and the delay of semantic correction. For example, when the above-mentioned simultaneous interpretation scheme is used, the delay of ASR can be 1.5 seconds, translation needs 3 seconds, and semantic correction needs 1.5 seconds. Therefore, after the user speaking (i.e., the sending end) finishes speaking, in addition to the transmission delay of the communication between the two ends, the user of the receiving end needs to spend an additional 6 seconds to hear the voice. Similarly, when the receiving party finishes listening to the above-mentioned voice, the user needs to experience similar delay when replying (i.e., when the receiving party becomes the sending party). It can be seen that although the simultaneous interpretation scheme can alleviate the problem of cross-language communication to a certain extent, in the scenarios of international conferences, remote work, and instant messaging that require quick response and real-time communication, the scheme still has the problem of delay and cannot meet the user's requirement for fast and real-time language conversion service.

[0069] In view of this, the embodiments of the present application provide a communication method, which is applied to a communication system including a sending end electronic device and a receiving end electronic device. The electronic device of the sending end can pre-receive user input, obtain and store corresponding translated voice based on the user input. After detecting the sending operation of the user, the electronic device of the sending end can send the pre-stored translated voice to the receiving end electronic device for playing. In some embodiments, the obtained translated voice can be stored in the electronic device of the sending end or other electronic devices that establish near-field connection (such as Bluetooth, WIFI, etc.) with the sending end based on the user input. For example, the user input and the obtained translated voice can be stored in a stylus or an audio acquisition device that establishes near-field connection with the sending end, so that the storage space of the near-field connected electronic device can be used to store the voice, reducing the storage pressure of the sending end electronic device. In some other embodiments, the obtained translated voice can also be stored in the server end or the cloud based on the user input, so that the sending can be faster and more efficient.

[0070] By implementing the method provided by the above-mentioned embodiments, the electronic device of the sending end can directly send the prepared translated voice when the user needs to send the voice, thereby greatly reducing the delay in the communication process, i.e., reducing the online waiting time of the user, and further providing a smooth and real-time online communication experience for the user.

[0071] The communication system includes a first electronic device and a second electronic device. In the following description, the first electronic device is taken as the sending-end electronic device, and the second electronic device is taken as the receiving-end electronic device.

[0072] It can be understood that, a prerequisite for the electronic device to perform the communication method is that the first user and the second user establish a first session in the first communication application. In some embodiments, the first user establishes the first session with the second user after starting the first communication application. The first user is a user using the first electronic device, and the second user is a user using the second electronic device. Optionally, after the first electronic device at the sending end starts the first communication application, the first user can log in to the first communication application through the first electronic device and establish a session with the second user. The present application does not limit the order of the first user logging in to the first communication application and starting the first communication application. Taking the first communication application as an online conference application as an example, a prerequisite for the communication system to perform the communication method is that the first user and the second user are participating in the same conference (i.e., during the existence of the first session), so that the communication demand is generated.

[0073] Under the above premise, the electronic device performing the communication method can distinguish between uplink processing and downlink processing. The uplink processing refers to the processing of the sending-end electronic device sending voice data, and the downlink processing refers to the processing of the receiving-end electronic device receiving voice data. The following describes the communication method provided by the present application from the uplink processing.

[0074] The first electronic device at the sending end has a first interface displayed thereon, and the first electronic device can perform the communication method by detecting user operations of a user in the first interface. FIG. 4 is a schematic diagram of a functional area distribution of a first interface according to an embodiment of the present application. In some embodiments, when the first electronic device displays the interface shown in FIG. 2 or FIG. 3, the first electronic device detects an operation of a user on a control, for example, a “translation” control (not shown in the figure), and displays the user interface shown in any one of FIG. 4, FIG. 5, or FIG. 6.

[0075] As shown in FIG. 4, the first interface can be divided into a plurality of functional areas. In FIG. 4, from top to bottom, there are a mode switching area, a shortcut voice area, a conversation voice storage area, and a user input area. In some embodiments, the electronic device can include at least the conversation voice storage area and an input box. The input box can be in the user input area or in the shortcut voice area.

[0076] The mode switching area has the following three voice sending modes for the user to select: sending original voice, sending translated voice, and sending both original voice and translated voice. It can be understood that the first electronic device as the sending end can detect the voice sending mode selected by the user in real time during the use of the first communication application by the user, and then send different voice data to the electronic device (i.e., the second electronic device) of the receiving end after detecting the user operation of sending voice. In other words, during the execution of the communication method provided by the embodiment of the application, the user can switch the voice sending mode at any time.

[0077] The original voice sending mode refers to that the first electronic device sends voice data of the language used by the sending end to the second electronic device. It can be understood that one or more texts are displayed in the first interface, for example, the text content of "Hello, I have a few questions to ask" in FIG. 4 is a text, and the voice data of the language used by the sending end is actually the voice data of the original text. Optionally, the voice data of the original text of the text can be obtained by the microphone of the first electronic device. Optionally, the voice data of the original text of the text can be generated by the first electronic device based on the original text.

[0078] The translated voice sending mode refers to that the first electronic device does not directly send the voice data of the original text, but sends the voice data of the translation of the text to the second electronic device after replacement. The replaced voice data of the translation of the text can be obtained by translation conversion based on the voice data of the original text. For example, the voice data of the original text is voice in language 1, and the translated voice data is voice in language 2. Optionally, the first electronic device can generate voice data of the translation of the text similar to the timbre and tone of the user based on the voiceprint cloning technology, so that the generated voice data of the translation of the text sounds more like the voice of a real person, thereby increasing the comfort and acceptance of the user of the receiving end when listening to the voice broadcast. It can be understood that the first electronic device can directly execute the voiceprint cloning step based on the voice data of the user using language 1 picked up by the microphone, without having to pre-record voice in language 2 for voiceprint cloning.

[0079] The original-translation dual mode refers to that the first electronic device sends not only the voice data of the original text but also the voice data of the translation of the text when sending the voice. For example, the first electronic device can send the voice data of the original text and then send the voice data of the translation of the text, so as to take into account the case that both personnel using the original language (such as language 1) and personnel using the translation language (such as language 2) exist in the receiver of online communication. In addition, by splicing the voice data of the original text and the voice data of the translation of a piece of text to realize the sending and playing of continuous audio stream, the delay of voice playing in the process of online communication can be effectively reduced, and thus the online user communication tends to be smooth.

[0080] The shortcut voice area can display some commonly used text and / or voice identifiers pre-stored in the electronic device. Taking the voice translation mode as an example, the pre-stored voice includes voice text and voice data of the translation of the text. The pre-stored commonly used voice is also referred to as a shortcut voice, and the text of the shortcut voice can be displayed in the first interface in the form of a sentence control, for example, the box in which “Hello, I have a few questions to ask” in FIG. 4 is a sentence control. In this way, the user can select the voice data to be sent according to the text in the subsequent process, without the need to click to play the voice to determine the voice data to be sent. After detecting a user operation acting on the sentence control in the shortcut voice area, the first electronic device can send the voice data of the translation of the text on the sentence control to the second electronic device. It can be understood that the user operations acting on various controls appearing in the present text include but are not limited to click operations on the controls. In addition to the click operation, the user operation acting on the sentence control in the shortcut voice area can also include dragging the sentence control in the shortcut voice area to the conversation voice storage area or the user input area. Optionally, the shortcut voice can include commonly used expressions in online communication, and the commonly used expressions can be pre-stored in the first electronic device. Optionally, the shortcut voice can be summarized by the first electronic device based on the user in the previous call or meeting. The summarizing step can be implemented by a large language model pre-stored in the first electronic device. Optionally, the shortcut voice can be obtained by the first electronic device from the input box of the shortcut voice area. For example, after the user inputs text in the input box of the shortcut voice area, the electronic device obtains the voice data corresponding to the text. Or the user inputs voice in the input box of the shortcut voice area. Optionally, in order to facilitate the user to select the voice data to be sent according to the text in the subsequent process, without the need to click to play the voice to determine the voice data to be sent, the voice input by the user can be converted into text and displayed in the shortcut voice area. In some embodiments, the user can select to import content (such as text, audio) from the memo as user input. It can be understood that when the first electronic device detects that there is content in the input box in the shortcut voice area and detects a user operation acting on the “custom shortcut voice” control as shown in FIG. 4, the first electronic device can generate a corresponding shortcut voice. Optionally, the shortcut voice can be obtained by the first electronic device from the input box of the user input area. For example, after the user inputs voice data or text in the user input area, the user clicks the control (such as the “prestore” control, not shown in the figure) of the user input area, pre-stores the voice data corresponding to the voice data or the text in the electronic device, and displays the text and / or voice identifier in the shortcut voice area.

[0081] In addition, the first electronic device can display a play control with the text of the shortcut voice. Upon detecting a user operation on the play control, the first electronic device can play voice data of a translation of the text. The first electronic device can also display a delete control with the text of the shortcut voice. Upon detecting a user operation on the delete control, the first electronic device can remove the shortcut voice from the shortcut voice area and delete voice data of a translation of the text of the shortcut voice pre-stored in the first electronic device. In some embodiments, the voice identifier can include the play control. In some embodiments, the play control can be text corresponding to the voice, and a user operation on the text, such as a long press, can play voice data corresponding to the text.

[0082] Similarly, in the original voice mode, the pre-stored voice includes text of the voice and voice data of the original text of the text. In the original-translation dual voice mode, the pre-stored voice includes text of the voice and spliced voice data of the text, which is spliced from voice data of the original text of the text and voice data of a translation of the text. As shown in FIG. 4, the text of the voice can be the original text. In these two modes, the controls in the shortcut voice area are similar to the controls in the translation voice mode, and the relevant descriptions above can be referred to and will not be repeated here.

[0083] The user input area can include an input box. The input box is configured to receive the text input by the first user, i.e., the first user input. The input by the first user can include, but is not limited to, keyboard input, voice input, and the like. In addition, as shown in FIG. 4, the first electronic device can display a "more" control (i.e., the control including a plus sign in the input box in FIG. 4) in the input box. After detecting a user operation on the "more" control, the first electronic device can provide more user input methods, such as text import, and the like. For example, the user can select to import text content from a memo as the user input. In some embodiments, the first electronic device can perform segmentation processing on the imported long text to divide the long text into multiple segments of the text, thereby helping the second user to better understand and digest the information. Alternatively, the first electronic device can perform segmentation processing on the long text based on segmentation of the long text itself. Alternatively, the first electronic device can perform segmentation processing on the long text based on natural language processing (NLP) technology. The embodiments of the present application do not make special limitations on the specific input method used by the user and the technical means used by the first electronic device for segmentation processing on the imported long text. The first electronic device can generate speech data of the translation of the text based on the text in the input box. For the text input by the user through the keyboard or through text import, the first electronic device does not have speech data of the original text of the text, and the first electronic device can also generate speech data of the original text of the text based on the text in the input box. Alternatively, in the above process, the first electronic device can use TTS technology to generate speech data of the text based on the text in the input box, and the speech data of the text includes one or more of the following: speech data of the original text of the text, speech data of the translation of the text. It can be understood that when the user inputs the text in the input box, the text input by the user and the generated speech data of the text can be temporarily stored in a third storage space in the first electronic device. The third storage space can be a temporary cache space matched with the input box, and is configured to temporarily cache the text and the speech data of the text before the user triggers the pause control or the send control.

[0084] The first electronic device can also display a temporary storage control in the user input area. Optionally, the first electronic device can display the temporary storage control only when there is user input text in the input box. After receiving the user input text, upon detecting a user operation of pausing sending the text (detecting a user operation acting on the temporary storage control), the first electronic device can store the voice data of the text in the input box as temporarily stored voice to be sent, and display the text in the conversation voice storage area. Specifically, after triggering the temporary storage control, the first electronic device can store the voice data of the text in the input box in a second storage space. This second storage space is a storage space provided by the communication method of the present application. When the user triggers the sending of voice data, the voice data stored in the second storage space can be transferred and stored in the first storage space.

[0085] In addition to the above-mentioned temporary storage control, the first electronic device can also display a sending control in the user input area. The sending control displayed in the user input area is also called the first sending control. Similarly, the first electronic device can display the first sending control only when there is user input text in the input box. After detecting a user operation of sending the text (such as detecting a user operation acting on the first sending control), the first electronic device can store the voice data of the text in the input box as voice to be sent, and display the text in the conversation voice storage area. Specifically, after triggering the first sending control, the first electronic device can store the voice data of the text in the input box in the first storage space, which can be accessed by the first communication application. It can be understood that the processes of executing the communication method provided by the present application and the first communication application can complete inter-process communication (IPC) through shared storage. Therefore, the first storage space is the shared storage. After triggering the first sending control, the first electronic device transfers and stores the voice data of the text in the input box from the third storage space to the first storage space, so that the first communication application can access the first storage space to obtain the voice data, and transmit the voice data to the first communication application of the receiving end.

[0086] After the first sending control is triggered, if no voice is being sent at present, the voice will be sent directly by the first electronic device to the second electronic device. If a voice is being sent at present, the first electronic device can add the voice to a sending queue and play the voice in the sending queue according to the order. That is, when the electronic device detects that the user edits multiple voices and the user sends the multiple voices, the electronic device can add the multiple voices to the sending queue according to the sending order. It can be understood that the first storage space has a limited storage capacity, for example, can store only 20 ms of voice data at a time. Therefore, only a part of data at the front of the sending queue enters the first storage space. Therefore, when the first electronic device detects the user operation of sending multiple texts continuously, the first electronic device can store the voice data of the multiple texts in the first storage space according to the sending order of the multiple texts. The first communication application program can access the first storage space and take away the voice data when the voice data exists, and then transmit the voice data to the first communication application program of the second electronic device and then play the voice data. It can be understood that after the first communication application program takes away the voice data from the first storage space, the first storage space no longer stores the voice data. Therefore, when the voice data of the text of the sentence control is played or the delete control beside the sentence control is triggered by the user, the voice data of the text of the sentence control leaves the sending queue. Optionally, the first electronic device can mark the sending order beside the multiple texts in the first interface to help the user determine whether the current sending order is suitable. If not, the user can choose to cancel sending or pause sending.

[0087] The conversation voice storage area is used to display at least one of the text of the voice that has been sent, the text of the voice that is being sent, and the text of the voice that is temporarily stored for sending. The text of the voice can also be displayed in the form of a sentence control on the first interface. As shown in FIG. 4, optionally, the user name, date, and time of sending the voice can be displayed above each sentence control. In some embodiments, the conversation voice storage area can also be referred to as a conversation area. It can be understood that the reason why the voice that is temporarily stored for sending is temporarily stored in the device is that the user does not send the voice temporarily, for example, in a meeting, the user has not yet spoken. The reason why the shortcut voice is pre-stored in the device is that the frequency of using the voice in the conversation can be high, and the pre-stored voice data can facilitate the user to send the voice multiple times. In terms of technology, the text of the voice that is temporarily stored for sending is obtained by the first electronic device through the input box of the user input area, and the text of the shortcut voice can be obtained by the first electronic device through the input box of the shortcut voice area or pre-defined in the first electronic device by the developer.

[0088] The voice being sent can include all voice data in the sending queue. As can be appreciated, a piece of voice data is composed of multiple voice packets, and the first electronic device can determine the progress of the current voice playing according to the number of voice packets sent, and then distinguish and display the sent part and the unsent part in the voice being sent. The unsent part is also referred to as the first part voice data, and the sent part is also referred to as the second part voice data. For example, as shown in FIG. 4, the first electronic device can display the sent part (such as the phrase "first") in the voice being sent in bold. For the voice being sent, the first electronic device can display a pause control next to the sentence control of the piece of voice, so as to simulate the situation that the user can be interrupted by other people in real communication, and improve the user experience. After detecting a user operation of pausing sending the text (such as detecting a user operation acting on the pause control), the electronic device can stop sending the remaining unsent voice packets of the piece of voice. Specifically, after detecting the user operation acting on the pause control, the first electronic device can stop storing the first part voice data in the first storage space, and the second part voice data has been stored in the first storage space before the user operation of pausing the control. As can be appreciated, if the first electronic device detects more than one piece of voice being sent, that is, there are multiple pieces of voice in the sending queue, after detecting the user operation acting on the pause control, the first electronic device can stop sending all the remaining voice packets in the sending queue. In other words, when there is more than one piece of voice being sent, the user triggers the pause control next to one of the sentence controls, and the first electronic device will stop playing the remaining unsent part of the piece of voice and all the voice being sent after the sentence in the sending queue. Specifically, the first electronic device can realize the function of pausing by emptying the sending queue. Optionally, after detecting the user operation acting on the pause control, the multiple pieces of voice in the sending queue that have not started sending voice packets can be converted into temporarily stored voice to be sent, and displayed in the conversation voice storage area.

[0089] The first electronic device can display a sending control, which is also referred to as a second sending control, along with the text of the temporarily stored voice to be sent. When a user operation of sending the text is detected (e.g., when a user operation on the second sending control is detected), the first electronic device can add the voice to a sending queue, and the voice is changed from a temporarily stored voice to a voice being sent. Optionally, the temporarily stored voice to be sent can be temporarily stored in the first electronic device in the form of a queue, which is also referred to as a temporary storage queue. When a user operation on the second sending control is detected, the first electronic device can add the voice and voices after the voice in the temporary storage queue to the sending queue in sequence, and then send the voice to the second electronic device, so that the user does not have to perform multiple operations on the second sending control, and the user operation is simplified. In some embodiments, after the user triggers the second sending control, the first electronic device transfers the voice data from the second storage space to the first storage space. It can be understood that, because the temporarily stored voice to be sent and the shortcut voice are different, the first electronic device can optionally store the two types of voice in different storage areas of the second storage space, for example, after the temporarily stored voice to be sent is obtained through the input box of the user input area, the first electronic device can store the temporarily stored voice to be sent in a first storage area of the second storage space; after the shortcut voice is obtained through the input box of the shortcut voice area, the first electronic device can store the shortcut voice in a second storage area of the second storage space. Thus, after a user operation of sending the text is detected, the first electronic device can obtain corresponding voice data from different storage areas of the second storage space to send to the second electronic device. The first electronic device can also distinguish the temporarily stored voice to be sent and the shortcut voice in the second storage space by using a flag bit or other special identifier, and the present embodiments do not specially limit how the first electronic device distinguishes the temporarily stored voice to be sent and the shortcut voice in the second storage space.

[0090] Optionally, a modification control can be provided beside the sentence control of the voice being sent and the voice to be sent. After the first electronic device detects a user operation of modifying the text (e.g., detecting a user operation acting on the modification control), the first electronic device can copy the text in the sentence control to the input box of the user input area, so that the first electronic device can receive the modified text, obtain and store the voice data of the translation of the modified text, and send the voice data of the translation of the modified text to the second electronic device or store it in the local (i.e., in the second storage space) after the user selects “send” or “store” again. In addition, after detecting the user operation of modifying the text, the first electronic device also transfers the voice data of the translation of the text from the first storage space or the second storage space to the third storage space. Optionally, a modification control can also be provided beside the sentence control of the voice being sent and the voice to be sent. It can be understood that the user operation acting on the modification control beside the voice being sent indicates that there may be an error in the expression of the previously sent voice. Therefore, after detecting the user operation acting on the modification control beside the voice being sent, the first electronic device can not only copy the text in the sentence control to the input box below, but also supplement a conjunction such as “sorry, please correct” before the text. Similarly, after the text is copied to the input box below, the voice data corresponding to the text is also stored in the third storage space. Optionally, the conjunction can be pre-set in the first electronic device by the developer and randomly selected by the first electronic device after the modification control is clicked each time. Optionally, the conjunction can be generated based on NLP technology. Compared with the pre-set solution, the conjunction generated based on NLP technology is more fluent in sentence expression. In addition, the first electronic device can also display a play control and a delete control with the text of each voice in the conversation voice storage area, which have the same functions as those in the shortcut voice area and will not be described again.

[0091] It can be understood that the first electronic device can display the sentence controls of different types of voices in different ways in the conversation voice storage area. Optionally, the first electronic device can display different colors for the sentence controls of different types of voices. For example, the first electronic device can display the sentence control of the voice being sent in orange, the sentence control of the voice to be sent in green, and the sentence control of the voice being sent in gray, i.e., the three sentence controls in the conversation voice storage area in FIG. 4 can be displayed in gray, orange, and green from top to bottom, so that the user can distinguish the types of the sentence controls.

[0092] Not limited to different color markings, the embodiments of the present application do not have special restrictions on how the electronic device distinguishes the sentence controls of different types of voices in the conversation voice storage area.

[0093] Not limited to the area distribution shown in FIG. 4, FIG. 5 is another distribution of functional areas of a first interface provided by an embodiment of the present application. As shown in FIG. 5, the mode switching area and the user input area are the same as the distribution of functional areas shown in FIG. 4. However, in the distribution of functional areas shown in FIG. 5, the first electronic device can display the sentence controls of the voice temporarily stored in the voice temporary storage area. In addition, the first electronic device can display the sent sentences and the sentences being sent in the voice sending area as shown in FIG. 5. It can be understood that when the first electronic device detects the input of the user in the user input area, if the first electronic device detects the user operation on the temporary storage control, the first electronic device can display the generated sentence control in the voice temporary storage area; if the first electronic device detects the user operation on the sending control, the first electronic device can display the generated sentence control in the voice sending area, so that the first electronic device can distinguish and display the types of voice more clearly by dividing the functional areas. Similarly, when detecting the user operation of continuously sending multiple voices, the multiple voices can be added to the sending queue by the first electronic device and sent to the second electronic device in sequence. Optionally, the first electronic device can store the voice temporarily stored in the sending queue in sequence, and then when detecting the user operation on the sending control, the first electronic device can add the voice temporarily stored corresponding to the sending control and the voice after the voice in the sending queue to the sending queue and display them in the voice sending area. For example, the voices of the sentence controls in the voice temporary storage area of the interface shown in FIG. 5 are stored in the sending queue in sequence, so that when detecting the user operation on the sentence control of "I'm sorry, I didn't quite understand, please repeat", the sentence and the next sentence "Second, we can refer to scheme A in the implementation scheme" can be sent to the second electronic device by the first electronic device. Optionally, the first electronic device can play the voice being sent according to the sending progress. The voice sending area is also referred to as the first area of the first interface, and the voice temporary storage area is also referred to as the second area of the first interface. It can be understood that for the distribution of functional areas shown in FIG. 4, the first area and the second area in the first interface partially overlap.

[0094] In the function area distribution diagram of the first interface shown in FIG. 4 or FIG. 5, the first electronic device can display only the language used by the sending-end user (i.e., the first user) in the sentence control (e.g., display only Chinese in the sentence control). For example, when the sending-end user uses Chinese, it can be understood that the user can also modify only the Chinese input in the input box when performing sentence modification, and then generate corresponding transliterated voice data based on the modified user input. However, in some other embodiments, as shown in the function area distribution diagram of the first interface in FIG. 6, the first electronic device can also display the translation in the sentence control. Specifically, the first interface already includes one or more texts, which include not only the original text but also the translation of the text. Specifically, when the user inputs the text, the first electronic device can display the translation of the text, so that when modifying the sentence in the input box, the user can modify not only the original text but also the translation, so that the accuracy of the final translation is further improved.

[0095] The following describes the specific process of uplink processing in different sending modes in the communication method provided in the embodiments of the present application, taking the case that the sending end uses Chinese and the receiving end uses English, in combination with the function area distribution diagram of the first interface. It can be understood that the uplink processing process of the first electronic device in different sending modes is ultimately to enable the first communication application program in the first electronic device to obtain different voice data from the first storage space, so that the first communication application program can send the different voice data to the first communication application program on the second electronic device, so that the voice can be played on the second electronic device. Therefore, the solution of the uplink processing process is described until the first communication application program on the first electronic device obtains the voice data.

[0096] FIG. 7 is a schematic diagram of an uplink processing process in a voice sending mode according to an embodiment of the present application. As shown in FIG. 7, when the first electronic device detects that the user selects the voice sending mode in the mode switching area, the first electronic device can pick up the voice of the user through the microphone (i.e., obtain the voice data of the original voice of the user), and process the picked-up voice by using an audio algorithm. It can be understood that the first electronic device can pick up the voice of the user through the microphone, i.e., the first electronic device receives the voice input of the user through the user input area shown in FIG. 4 to FIG. 6. The audio algorithm includes but is not limited to a noise reduction processing algorithm, an echo cancellation algorithm, and an automatic gain control algorithm. It can be understood that the microphone and the audio algorithm can be provided by the operating system of the first electronic device. Subsequently, the first electronic device can send the processed voice data to the first communication application program, so that the first communication application program can obtain the voice data, i.e., the voice data of the original text of the text input by the voice.

[0097] Fig. 8 is a schematic diagram of an uplink processing flow in a speech translation mode according to an embodiment of the present application. As shown in Fig. 8, when the first electronic device detects that the user selects the speech translation mode in the mode switching area of the interface shown in Figs. 4-6, the first electronic device can also pick up the user's speech through the microphone and process the picked-up speech by using the audio algorithm. Subsequently, the first electronic device can obtain the Chinese text corresponding to the processed speech by using the ASR technology, and display the Chinese text in the input box of the user input area shown in Figs. 4-6. Optionally, the first electronic device can obtain the user's correction to the ASR result from the input box. In one implementation, the first electronic device can translate the content in the input box to obtain the corresponding English text (i.e., the translation). In another implementation, the first electronic device can directly translate the English text based on the processed speech data by using the audio algorithm. It can be understood that, if the corresponding translation is obtained based on the content in the input box, the first electronic device can receive the user's correction to the text in the input box, thereby improving the accuracy of the ASR result and further improving the accuracy of the translation. After obtaining the translation, the first electronic device can display the translation as shown in Fig. 6. Optionally, the first electronic device can not display the translation as shown in Fig. 4 or Fig. 5. Based on the translation, the first electronic device can generate the speech data (here, the English speech) of the translation of the text by using the TTS technology, and send the English speech to the third storage space, i.e., the speech data buffer splicing module shown in Fig. 8 and Fig. 9. When the user triggers the temporary storage control, the above-mentioned English speech is transferred from the third storage space to the second storage space (not shown in Fig. 8 and Fig. 9). When the user triggers the second sending control, the first electronic device can transfer the above-mentioned English speech from the second storage space to the first storage space according to the sending mode of the speech translation. Without being limited to the scheme of temporary storage and then sending, when the user triggers the first sending control, the first electronic device can also transfer the above-mentioned English speech from the third storage space to the first storage space according to the sending mode of the speech translation. After being sent to the first storage space, the English speech can be obtained by the first communication application.

[0098] Compared with the uplink processing flow in the translation mode shown in FIG. 8, the uplink processing flow in the original-translation dual-transmission mode shown in FIG. 9 only has a difference in the voice data (i.e., the voice data stored in the first storage space) sent to the first communication application in the last step. Specifically, after detecting the user operation of sending the text, if it is detected that the user-selected sending mode is the original-translation dual-transmission mode, the first electronic device can perform a splicing operation on the Chinese voice and the English voice in the third storage space or the second storage space to obtain spliced voice data, for example, by concatenating the English voice after the Chinese voice, and then sending the spliced voice data to the second electronic device. In addition, the other steps in the uplink processing flow in the original-translation dual-transmission mode are the same as those in FIG. 8, and the related description above can be referred to, and will not be described here.

[0099] Preferably, regardless of the sending mode, after the user triggers the temporary storage control, the first electronic device sends the voice data of the original text of the text to the second storage space. This is because in actual application, the user can change the sending mode at will, that is, there can be a case that after the user triggers the temporary storage control, the user modifies the sending mode from the translation mode to the original-translation dual-transmission mode. If the first electronic device does not send the voice data of the original text of the text to the second storage space after the temporary storage control is triggered, the first electronic device cannot find the voice data of the original text of the text from the second storage space after the sending mode is switched, and thus cannot transfer and store the voice data of the original text of the text into the first storage space for the first communication application to obtain.

[0100] In the processes shown in FIGS. 8 and 9 above, the first electronic device can directly determine the language of the original voice according to the user input. In order to successfully perform the method, the first electronic device also needs to determine the language of the translated text in the translation process.

[0101] In some embodiments, the language of the translated text can be set by the user in the first interface in advance, and then the first electronic device can directly perform the translation operation based on the setting.

[0102] Optionally, the language of the translated text can be set by the user holding the first electronic device in the first interface in advance. It can be understood that for a meeting attended by people from multiple countries, the language used in the meeting can be determined in advance. For example, for a meeting attended by people from Germany, the United Kingdom, and China, the attendees can determine in advance to communicate in English. In this case, the first electronic device can directly determine the language of the translated text in the translation step according to the selection operation of the holder on the first electronic device.

[0103] Optionally, the language of the translated text can be set by the second user of the second electronic device and then notified to the first electronic device. For example, a language selection control can be provided on the display interface of the second electronic device. If the second user is a German, upon detecting that the user selects German as the language, the second electronic device can generate a first notification and send the first notification to the first electronic device. The first notification is used to notify the first electronic device that the second user needs to receive the translated text in German. Upon receiving the first notification, the first electronic device can set the language of the translated text in the translation step to German according to the first notification.

[0104] It can be understood that when there are multiple receiving ends (i.e., the second electronic device is multiple electronic devices) in the online conversation, the first electronic device can generate voice data of multiple different translated texts of the text based on multiple first notifications sent by the second electronic device, for example, both German voice data and English voice data, store the generated voice data of multiple different translated texts of the text in the first storage space, and then forward the voice data of multiple different translated texts of the text to different second electronic devices by the first communication application, for example, forward the German voice data to the second electronic device held by the German. It can be understood that the first notification can include the correspondence between the language and the electronic device, thereby helping the first electronic device to determine which voice data of the translated text to send to which second electronic device.

[0105] In some embodiments, the first electronic device can perform language recognition on the received voice data, thereby determining the language of the translated voice based on the recognition result. Specifically, the first electronic device can obtain the voice data sent by the second electronic device from the first communication application. Then, the first electronic device can use a model such as a convolutional neural network, a recurrent neural network, or a long short-term memory network to recognize the language of the voice data, thereby determining the language of the translated voice. For example, the first electronic device recognizes that the language used by the second user in the meeting is German based on the model, and then the first electronic device can set the language of the translated text in the translation operation to German.

[0106] It can be understood that if the first electronic device has generated a temporarily stored voice to be sent or has added the voice to the sending queue before the setting, that is, the first electronic device has stored the voice data into the first storage space or the second storage space, the first electronic device can re-execute the process as shown in FIG. 8 or FIG. 9 to generate the voice data of the new translation of the text and replace the voice data of the old translation of the text. In the above example, the first electronic device can be set to default translation into English. If the user has temporarily stored a voice before switching the translation into German, the temporarily stored voice to be sent is an English voice. When the first electronic device detects that the language used by the second user is German, the first electronic device can first automatically switch the language of the translation in the translation step to German, and then the first electronic device can determine whether there is a temporarily stored voice to be sent or whether there is a voice in the sending queue that has not been sent, that is, the first electronic device can determine whether the second storage space is empty at present. In this step, the first electronic device can determine that there is a temporarily stored voice to be sent, and then the first electronic device can re-execute the translation step on the temporarily stored voice to be sent to obtain a German text, and obtain the German voice corresponding to the German text through TTS voice broadcast. Then, the first electronic device can replace the English voice of the same text stored in the second storage space with the German voice. After the above steps, the audio data sent by the first electronic device to the second electronic device through the first communication application can be switched to German.

[0107] The above FIG. 7 to FIG. 9 are used to explain the overall process of the uplink processing by taking the voice input as an example. It can be understood that the user input is not limited to the voice input, if the user selects to directly input in the input box by keyboard or text import and the like, the first electronic device can directly execute the translation, TTS broadcast and the like based on the user input to obtain and store the voice data of the translation of the text, without executing the audio algorithm and ASR step. In addition, in the original sound translation sound double sending mode, the first electronic device can also directly execute the TTS broadcast step based on the text input by the user to generate the corresponding Chinese voice as the voice data of the original text of the text for subsequent splicing process. The remaining steps and detailed process can be referred to the related description above, which will not be described here.

[0108] In actual use, the above communication function can exist in the electronic device as an independent application, or can be embedded into the first communication application as a functional module. If the above communication function is embedded into the first communication application as a functional module, the first communication application should provide a control to jump to the first interface. When the first electronic device detects that the jump control is clicked, the first electronic device can enter the first interface from other use interfaces of the first communication application.

[0109] If the communication function can exist as an independent application in the electronic device, the first interface can be entered through the floating ball as shown in FIG. 10. It can be understood that when the first electronic device detects the first operation, the first electronic device can jump from FIG. 10 to the floating window as shown in, for example, FIG. 11 or FIG. 12, and the floating window is used to display the first interface. The first operation includes but is not limited to the click operation of the user on the floating ball in FIG. 10. The communication method can run simultaneously with the first communication application on the electronic device through the floating window as shown in FIG. 11 or FIG. 12, to realize parallel operation.

[0110] Due to the size of the display interface of the electronic device, when displaying the first interface, the first electronic device can not display all the function areas in FIG. 4 or FIG. 5, but can only display part of the function areas. Taking the function area distribution shown in FIG. 5 as an example, the first electronic device can only display the mode switching area, the voice storage area, and the user input area as shown in FIG. 11, or can only display the mode switching area, the voice sending area, and the user input area as shown in FIG. 12. Optionally, after detecting the user operation on the sentence control in the voice storage area, or after detecting the user operation on the sending control in the user input area, the first electronic device can jump from the interface shown in FIG. 11 to the interface shown in FIG. 12, so that the limited display space can be reasonably utilized to provide the user with a clearer and more focused user experience. Optionally, the first electronic device can also set a jump control in the interface shown in FIG. 11, and the user operation on the jump control can also achieve the effect of the above interface jump.

[0111] Optionally, the electronic device can also display all the function areas as shown in FIG. 4 or FIG. 5 in the floating window. The embodiments of the present application do not specially limit which function areas are displayed in the first interface.

[0112] It can be understood that the user operation on the control is not limited to the user operation on the control, and all the above-mentioned storage controls, pause controls, sending controls, etc. The user operation on the control can also be replaced by a gesture operation on the text, for example, the user operation of sending the text can include dragging the sentence control from the shortcut voice area to the user input area. For another example, the user operation of pausing the sending of the text can include double-clicking the sentence control being sent. For another example, the user operation of storing the text can include long pressing the text in the input box. The gesture operation on the text can be pre-set by the developer and stored in the electronic device, and the embodiments of the present application do not specially limit whether the user operation of storing, pausing, sending the text is realized by the control or by the gesture operation on the text and other ways.

[0113] In some embodiments, the first interface further includes the following function: the first electronic device can merge speech and display the merged statement control. Specifically, after detecting a merging operation (e.g., pinching) between the first text and the second text in the first interface, the first electronic device can concatenate the first text and the second text into a third text. The speech data of the translation of the third text is obtained by concatenating the speech data of the translation of the first text and the speech data of the translation of the second text. The merging operation includes a user operation of dragging the first text to the second text. For example, as shown in Figure 13, after long-pressing the "Please explain again" statement control, the user can drag the statement control to the "Sorry, I didn't quite understand" statement control. After detecting the above operation, the first electronic device can merge the two statements, including merging the two statement controls and merging the speech data of the two statements. Specifically, the first electronic device can append the dragged statement after the statement at the pause position, so that the electronic device can display as shown in Figure 14, that is, the merged statement control in the above example is displayed as "Sorry, I didn't quite understand, please explain again". After merging the voice data of the two statements, the first electronic device can use the merged voice data to replace the two independent voice data of the two statements in the second storage control.

[0114] The aforementioned voice merging function design allows users to easily reassemble voice information, improving the flexibility of voice information processing and the convenience of user operation, thereby enhancing the user's online communication experience.

[0115] The first electronic device can also not perform the merging processing on the plurality of speeches, but only stack display the sentence controls of the plurality of speeches. The above-mentioned stack display of the sentence controls refers to that the first electronic device displays a plurality of sentence controls in the first interface through a first display area. Specifically, the first display area displays only one sentence control at a time; when it is detected that a second operation is performed by the user on the display area (such as a downward sliding operation performed in the first display area or a click on a switching button beside the first display area, etc.), the first electronic device can switch the sentence control displayed in the display area, thereby improving the use efficiency of the display area in the first interface. Optionally, the first display area can display a play control. When a user operation acting on the play control is detected, the first electronic device can play the speech corresponding to the sentence control currently displayed in the first display area. Optionally, after the first time of the user operation acting on the play control, the first electronic device no longer plays the speech corresponding to the sentence control currently displayed in the first display area. Optionally, within the first time after the previous speech is played, if the first electronic device does not detect the second operation of the user, the first electronic device no longer plays the speech corresponding to the sentence control currently displayed in the first display area. In some embodiments, for each sentence control, the first electronic device can play the speech corresponding to the sentence control only once when the second operation is not detected. Thus, through the second operation and the function of stack display, the user does not need to click the play control every time to listen to the speech corresponding to the sentence control, thereby simplifying the user operation.

[0116] Based on the above scheme, the first electronic device can pre-acquire the speech data of the translation of the text to be sent, and immediately transmit the speech data of the translation of the text after detecting the user operation of sending the text, thereby providing a more smooth online communication experience for the user. Moreover, since the user can pre-store the speech data, the user can continue to think about the subsequent expression content when sending the speech, thereby significantly optimizing the communication efficiency and expression coherence of the user.

[0117] It can be understood that, in the online communication process, the first user can change from the role of the electronic device of the sending end to the electronic device of the receiving end after finishing speaking, and similarly, the second user can also become the user who speaks at any time, that is, the second electronic device can also change from the electronic device of the receiving end to the electronic device of the sending end. If the second electronic device can also perform the above-mentioned communication function, in order to meet the demand of the user to speak in time and provide a smooth online communication experience, the second electronic device, as the electronic device of the receiving end, can also display the same first interface as the first electronic device before the device role conversion, thereby facilitating the second electronic device to change the device role at any time, that is, facilitating the second user to become the speaker in the online communication at any time.

[0118] It can be understood that the above scheme has the benefit of processing voice data mainly in that the user at the receiving end obtains a better listening experience. However, for the first user, when the first user finishes speaking, the device role changes, and the first electronic device changes from the electronic device at the sending end to the electronic device at the receiving end. At this time, if the electronic device at the sending end (i.e., the second electronic device) cannot perform the above communication method, and the above communication method only includes the uplink processing flow, even if the first electronic device of the first user can perform the above communication method, the first user's receiving experience will not be significantly optimized when receiving the voice of the second user. Therefore, optionally, the above communication method can also include a downlink processing flow.

[0119] FIG. 15 is a flow diagram of a downlink processing according to an embodiment of the present application. As shown in FIG. 15, after receiving the voice data, the first electronic device can play the voice data through a speaker or a headset as shown in FIG. 15. At the same time, the first electronic device can obtain the text corresponding to the voice through the ASR technology, and then translate it into Chinese and present it on the display interface. However, this scheme still requires the user to read the subtitles and understand their meaning. This means that when participating in online communication, the user still needs to divide attention to the text on the screen in order to keep up with the progress of the conversation. Therefore, although this method provides a means of auxiliary understanding, the user may need to constantly switch attention between listening and reading, and cannot provide a natural and smooth online communication experience.

[0120] In view of this, the embodiment of the present application provides another downlink processing procedure as shown in FIG. 16. Specifically, the voice data received by the first electronic device can also be referred to as voice data of the original text of the fourth text. After receiving the voice data of the original text of the fourth text, the first electronic device can first detect whether it has a dual-channel playback capability. Specifically, the dual-channel playback capability can be provided by a single audio playback device capable of channel separation (such as a Bluetooth headset), or can be provided by multiple audio playback devices connected to the first electronic device. The channel separation capability means that the audio playback device can output different audio data to the left and right channels respectively. If the first electronic device detects that it has the dual-channel playback capability, the first electronic device can play the received voice in one of the channels or one of the audio playback devices. At the same time, on the basis of the scheme shown in FIG. 15, the voice played in the other channel or the other audio playback device is subjected to ASR recognition and translation, and the translation result is subjected to TTS technology to obtain voice data (such as Chinese voice) of the translation of the fourth text. Subsequently, the first electronic device can play the voice data of the translation of the fourth text obtained based on the TTS operation through the other channel or the other audio playback device. For example, if the second electronic device transmits English voice (i.e., the original text of the fourth text is an English text), through the downlink processing procedure described in FIG. 16, in the Bluetooth headset connected to the first electronic device, the right ear can play the original English voice, and at the same time, the left ear can play the Chinese voice generated based on TTS, so that the user can choose to listen to the voice in which language, so that the user can participate more deeply in the conversation, and provide a better online communication experience for the user.

[0121] It can be understood that when the first electronic device and the second electronic device can both execute the communication method provided by the embodiment of the present application, and the sending mode set by the sending-end electronic device is "original sound and translated sound dual sending", if there is no additional processing, the user will hear two identical voice broadcasts in the left earphone of the Bluetooth headset, which does not conform to the user's communication habits.

[0122] In the above example, a preferred implementation is that when playing the voice, the first electronic device (i.e. the electronic device as the receiving end at this time) can detect whether there are two repeated voices in the broadcast content. Specifically, if the first electronic device receives the voice data of the translation of the fourth text sent by the second electronic device after receiving the voice data of the original text of the fourth text sent by the second electronic device, the first electronic device determines not to play the voice data of the translation of the fourth text sent by the second electronic device. It can be understood that the same tone refers to the voice data of the translation of the fourth text sent by the second electronic device, which has the same tone as the voice data of the translation of the fourth text generated by the first electronic device based on the voice data of the original text of the fourth text received by the receiving end. Optionally, the first electronic device can detect the tone of the voice based on a neural network model, and when the tones of the above two voice data of the translation of the fourth text in the broadcast content are consistent, the first electronic device can determine that there are two repeated voices in the broadcast content.

[0123] Another optional implementation is to establish a communication mechanism between electronic devices. Specifically, when the second electronic device detects that the user selects the "original voice and translation voice double sending" sending mode, the second electronic device can send a second notification to the first electronic device through the above-mentioned first communication application, and the second notification is used to inform the electronic device as the sending end that the first electronic device adopts the "original voice and translation voice double sending" sending mode. After receiving the above-mentioned second notification, after receiving the voice data of the original text of the fourth text and the voice data of the translation of the fourth text, the first electronic device can only play once, i.e. determine not to play the voice data of the fourth text received afterwards.

[0124] It can be understood that the above scheme is based on the above example, and the first electronic device is taken as an example of the electronic device as the receiving end after the device identity conversion. However, when the first electronic device is as the electronic device as the sending end, as long as the condition of hearing the same voice broadcast from both ends is met, the second electronic device can also be as the electronic device as the receiving end. At this time, the second electronic device performs the same operation as the above-mentioned first electronic device, and can refer to the above related description, which will not be described again.

[0125] It can be understood that the first communication application described above is not limited to applications with high real-time requirements such as online meetings, but also includes social media and the like that can send voice. Therefore, the communication method provided by the embodiments of the present application can also include the following examples: the first electronic device can obtain a data packet of an English voice by performing the uplink processing procedure on the Chinese voice of the first user and transmitting the data packet to the second electronic device, but in the first communication application of the second electronic device, the English voice is not played, but is temporarily stored in the first communication application in the form of a voice bar or the like. That is, the second user holding the second electronic device can review the voice bar of the English voice sent by the first user at leisure. In the above example, the processing procedure of the communication method is the same as that of the communication method for playing voice data in real time, and reference can be made to the related description above, and details are not repeated.

[0126] FIG. 17 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. The electronic device can be the first electronic device described above, or the second electronic device described above.

[0127] As shown in FIG. 17, the electronic device can include a processor 211, a memory 212, a wireless communication processing module 213, a power switch 214, a display screen 215, an audio module 216, a speaker 217, and the like. The components in the electronic device are connected through a bus and communicate based on the bus.

[0128] The processor 211 can include one or more processing units, for example: the processor 211 can include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units can be independent devices, or can be integrated in one or more processors. The controller can generate operation control signals according to instruction operation codes and timing signals, and complete the control of fetching instructions and executing instructions.

[0129] The memory 212 is coupled to the processor 211 and is configured to store various software programs and / or sets of instructions. The memory 212 can be configured to store computer-executable program code, which includes instructions. The processor 211 is configured to perform various functional applications and data processing of the electronic device by executing the instructions stored in the memory 212. The processor 211 can also be configured to store instructions and data in a memory.

[0130] The memory 212 can include one or more random access memories (RAMs) and one or more non-volatile memories (NVMs). The random access memory can be directly readable and writable by the processor 211. The random access memory can be configured to store executable programs (e.g., machine instructions) of an operating system or other programs that are currently running, and can also be configured to store data of users and application programs, and the like. The non-volatile memory can also store executable programs and store data of users and application programs, and the like. The executable programs and user data stored in the non-volatile memory can be loaded in advance into the random access memory for direct reading and writing by the processor 211.

[0131] Executable program codes and data (including but not limited to various voice data) used to implement the voice processing method provided by the embodiments of the present application can be stored in the non-volatile memory. In the process of implementing the above communication method, the electronic device can load the executable program codes and user data of the non-volatile memory into the random access memory to provide a smooth online communication experience for the user.

[0132] The wireless communication processing module 213 can provide wireless communication solutions including WLAN, such as Wi-Fi, Bluetooth communication, ZigBee communication, NFC communication, infrared communication, UWB communication, and the like. In the embodiments of the present application, the electronic device can send and receive various data to other electronic devices based on the wireless communication functions provided by the wireless communication processing module 213, including but not limited to original voice data, translated voice data, first notifications, and the like.

[0133] The power switch 214 can be configured to control the power supply of the electronic device, and in turn, power the processor 211, the memory 212, the wireless communication processing module 213, the display screen 215, the audio module 216, the speaker 217, and the like.

[0134] The display 215 can be configured to display. The display 215 can include a display panel. A touch sensor can be disposed in the display 215. The touch sensor can be configured to detect a touch operation acting on or near the touch sensor. The touch sensor can transmit the detected touch operation to the application processor to determine a touch event type. In turn, the electronic device can provide visual output related to the touch operation through the display 215. The electronic device can implement a display function through the GPU, the display 215, the touch sensor, the application processor, and the like. In embodiments of the present application, the electronic device can display the use interface of the first communication application and the interface of the communication system to the user through the display function provided by the GPU, the display 215, the touch sensor, the application processor, and the like, and switch between various interfaces based on the click operation on various controls. In addition, the electronic device can receive the input of the user based on the display function, such as receiving the text content imported from the memo, and can also correct the text in the input box, and the like.

[0135] The audio module 216 can be configured to convert a digital audio signal into an analog audio signal for output, and can also be configured to convert an analog audio input into a digital audio signal. The speaker 217 can be configured to convert a transmitted audio signal of the audio module 216 into a sound signal. The electronic device can implement an audio play function through the audio module 216, the speaker 217, and the like. In embodiments of the present application, after receiving the voice data transmitted by the sending end, the electronic device can call the audio module 216 and the speaker 217 to play the received original voice and translated voice. In some embodiments, the audio module 216 can further include a microphone configured to convert a sound signal into an electrical signal. In embodiments of the present application, the electronic device can call the microphone to pick up the voice of the user, generate a voice packet, and then use the voice packet for voice input.

[0136] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device. In other embodiments of the present application, the electronic device can include more or fewer components than illustrated, or combine certain components, or split certain components, or different component arrangements. The components can be implemented in hardware, software, or a combination of software and hardware.

[0137] Those skilled in the art should be aware that, in the above one or more examples, the functions described in the embodiments of the present application can be implemented in hardware, software, firmware or any combination thereof. When implemented in software, the functions can be stored in a computer readable medium or transmitted as one or more instructions or code on a computer readable medium. The computer readable medium includes computer storage medium and communication medium, and the communication medium includes any medium that facilitates transfer of a computer program from one place to another. The storage medium can be any available medium that can be accessed by a general purpose or special purpose computer. The embodiments of the present application also provide a computer program product including a computer program, which can implement the steps of the above various method embodiments when the computer program is run on a processor.

[0138] The above detailed description of the embodiments of the present application has further detailed the purposes, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above is only a specific implementation of the embodiments of the present application, and is not used to limit the protection scope of the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made on the basis of the technical solutions of the embodiments of the present application should be included in the protection scope of the embodiments of the present application.

Claims

1. A communication method characterized by comprising: The method applied to a first electronic device comprises: The first electronic device starts a first communication application, and establishes a first session in the first communication application between a first user using the first electronic device and a second user using a second electronic device; The first electronic device detects a first operation, displays a first interface, and stores voice data of a translation of the text in the first interface; During the existence of the first session, a user operation of sending the text is detected, and the first electronic device sends the voice data of the translation of the text to the second electronic device.

2. The method of claim 1, wherein, Before the first interface displays one or more texts and the first electronic device stores voice data of a translation of the text, the method further comprises: The first electronic device displays an input box in the first interface; The first electronic device receives a text input by a user from the input box; The first electronic device acquires voice data of a translation of the text.

3. The method of claim 1, wherein, Before the first interface displays one or more texts and the first electronic device stores voice data of a translation of the text, the method further comprises: The first electronic device displays an input box in the first interface; The first electronic device receives a voice input by a user from the input box; The first electronic device converts the voice input into a text; The first electronic device acquires voice data of a translation of the text.

4. The method according to any one of claims 1 to 3, characterized in that, The text comprises an original text and a translation.

5. The method according to any one of claims 1 to 4, characterized in that, The first electronic device stores voice data of a translation of the text, specifically comprising: The first electronic device stores spliced voice data of the text, the spliced voice data being spliced from voice data of the original text and voice data of the translation, the voice data of the original text being obtained by a microphone of the first electronic device or being generated by the first electronic device based on the original text.

6. The method according to any one of claims 1 to 5, characterized in that, Before the detection of the user operation of sending the text, the method further comprises: After receiving the text input by the user, a user operation of temporarily storing the text is detected, and the first electronic device stores voice data of a translation of the text.

7. The method according to any one of claims 1 to 6, characterized in that, After the detection of the user operation of sending the text, the method further comprises: A user operation of pausing the sending of the text is detected, and the first electronic device stops sending the first part of the voice data to the second electronic device, the voice data of the translation of the text being composed of the first part of the voice data and a second part of the voice data, the second part of the voice data having been sent to the second electronic device by the first electronic device before the user operates the pause control.

8. The method according to claims 1-7, characterized by, Before the detection of the user operation of sending the text, the method further comprises: A user operation of modifying the text is detected, and the first electronic device displays the text in an input box in the first interface; The first electronic device receives a modified text input by a user from the input box; The first electronic device acquires and stores voice data of a translation of the modified text; The first electronic device detects a user operation of sending the text, and sends speech data of a translation of the text to the second electronic device. The first electronic device detects a user operation of sending the modified text, and sends speech data of a translation of the modified text to the second electronic device.

9. The method according to any one of claims 1 to 8, characterized in that, The first electronic device detects a user operation of sending the text, and sends speech data of a translation of the text to the second electronic device. The first electronic device detects a user operation of continuously sending a plurality of texts. The first electronic device sends speech data of translations of the plurality of texts to the second electronic device in a sending order of the plurality of texts.

10. The method of claim 9, wherein, After the first electronic device detects the user operation of continuously sending the plurality of texts, the method further includes: The first electronic device labels a sending order of each of the plurality of texts.

11. The method according to any one of claims 1 to 10, characterized in that, The one or more texts displayed in the first interface include predefined texts.

12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: After detecting a merging operation of a first text in the first interface and a second text in the first interface, the first electronic device splices the first text and the second text into a third text, speech data of a translation of the third text is obtained by splicing speech data of a translation of the first text and speech data of a translation of the second text, the plurality of texts include the first text and the second text, and the merging operation includes a user operation of dragging the first text to the second text.

13. The method according to any one of claims 1 to 12, characterized in that, The method further includes: The first electronic device receives speech data of original text of a fourth text sent by the second electronic device, and plays the speech data of the original text of the fourth text sent by the second electronic device.

14. The method of claim 13, wherein, The method further includes: The first electronic device detects that it has a dual-channel playback capability. The first electronic device translates the original text of the fourth text to obtain a translation of the fourth text. The first electronic device plays speech data of the original text of the fourth text through one channel and plays speech data of the translation of the fourth text through another channel.

15. The method of claim 14, wherein, The method further includes: If, after receiving the speech data of the original text of the fourth text sent by the second electronic device, the first electronic device further receives speech data of a translation of the fourth text sent by the second electronic device with the same tone, the first electronic device determines not to play the speech data of the translation of the fourth text sent by the second electronic device.

16. A method of communication, comprising: Applied to the second electronic device, the method includes: The second electronic device receives speech data of original text of a text sent by the first electronic device, and plays the speech data of the original text of the text sent by the second electronic device. The second electronic device detects that it has a dual-channel playback capability. The second electronic device translates the original text of the text to obtain a translation of the text. The second electronic device plays speech data of the original text of the text through one channel and plays speech data of the translation of the text through another channel.

17. The method of claim 16, wherein, The method further includes: If, after receiving the voice data of the first electronic device sending the original text of the text, the second electronic device also receives the voice data of the first electronic device sending the translation of the text in the same tone, the second electronic device determines not to play the voice data of the first electronic device sending the translation of the text.

18. An electronic device, the electronic device being a first electronic device, comprising: comprising one or more processors and one or more memories; wherein the one or more memories are coupled with the one or more processors, and the one or more memories are configured to store computer program code, the computer program code comprising computer instructions that, when executed by the one or more processors, cause performance of the method of any one of claims 1-15.

19. An electronic device, the electronic device being a second electronic device, characterized by comprising one or more processors and one or more memories; wherein the one or more memories are coupled with the one or more processors, and the one or more memories are configured to store computer program code, the computer program code comprising computer instructions that, when executed by the one or more processors, cause performance of the method of any one of claims 16-17.

20. A communication system comprising a first electronic device and a second electronic device, the first electronic device being the electronic device of claim 18, and the second electronic device being the electronic device of claim 19.

21. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when running on an electronic device, causes performance of the method of any one of claims 1-15 or 16-17.

22. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method of any one of claims 1-15 or 16-17.

Citation Information

Patent Citations

  • Meeting implementation method, device, equipment and system and computer readable storage medium

    CN108076306A

  • Method, system and device for converting characters into voice in chat scene and storage medium

    CN113495766A

  • Dialect translator for a speech application environment extended for interactive text exchanges

    US20080147408A1

  • Text input method, electronic device, and system

    WO2022095820A1