Call method and apparatus, electronic device, and readable storage medium
By translating and sending translated voice messages during a call, the problem of low communication efficiency caused by language barriers between the two parties is solved, achieving more efficient communication.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- VIVO MOBILE COMM CO LTD
- Filing Date
- 2025-10-16
- Publication Date
- 2026-04-30
AI Technical Summary
When the two parties in a call do not understand each other's language, communication efficiency is low.
By translating the voice during a call, and sending and playing the translated voice, both parties in the call can directly understand each other's conversation.
It improves communication efficiency between the two parties in a call and avoids communication barriers.
Smart Images

Figure CN2025128161_30042026_PF_FP_ABST
Abstract
Description
Communication methods, devices, electronic devices and readable storage media
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411470212.6, filed in China on October 21, 2024, the entire contents of which are incorporated herein by reference. Technical Field
[0003] This application belongs to the field of communication technology, specifically relating to a communication method, apparatus, electronic device, and readable storage medium. Background Technology
[0004] Currently, users can capture voice messages through electronic devices, and then use these captured voice messages to communicate with other users. For example, when a user is talking to another user, the electronic device can capture the user's voice output, send that voice message to the other user's electronic device, and receive voice messages from the other user's electronic device, thus enabling the user to communicate with other users.
[0005] However, if the two parties in a call do not understand each other's language, there may be communication barriers, resulting in lower communication efficiency. Summary of the Invention
[0006] The purpose of this application is to provide a call method, device, electronic device, and readable storage medium that can improve the communication efficiency of both parties in a call.
[0007] In a first aspect, embodiments of this application provide a call method applied to an electronic device corresponding to a first caller. The call method includes: when the first caller and the second caller are in a call, performing at least one of the following: upon receiving a first voice message from the first caller, sending the first voice message and a translated voice message of the first voice message, or a translated voice message of the first voice message, to the second caller; upon receiving a second voice message from the second caller, playing the second voice message and a translated voice message of the second voice message, or a translated voice message of the second voice message.
[0008] Secondly, embodiments of this application provide a communication device applied to an electronic device corresponding to a first party in a call. The communication device includes a sending module and a playback module. The sending module is configured to, when the first party and a second party are in a call, send the first voice message of the first party and a translated version of the first voice message, or a translated version of the first voice message, to the second party upon receiving the first voice message from the first party. The playback module is configured to, upon receiving the second voice message from the second party, play the second voice message and a translated version of the second voice message, or a translated version of the second voice message.
[0009] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.
[0010] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.
[0011] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.
[0012] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.
[0013] In this embodiment, when a first party and a second party are in a conversation, at least one of the following is performed: upon receiving a first voice message from the first party, sending the first voice message and its translation, or the translation of the first voice message, to the second party; upon receiving a second voice message from the second party, playing the second voice message and its translation, or the translation of the second voice message. In this solution, because the second voice message can be translated and the translated voice message is output, the first party can directly and clearly understand the content of the second party's conversation based on the translated voice message; furthermore, sending the translated voice message to the second party allows the second party to clearly understand the content of the first party's conversation, avoiding communication barriers between the two parties and thus improving communication efficiency. Attached Figure Description
[0014] Figure 1 is a flowchart of a call method provided in an embodiment of this application;
[0015] Figure 2 is an example diagram of a caller interface provided in an embodiment of this application;
[0016] Figure 3 is one example of a call interface provided in an embodiment of this application;
[0017] Figure 4 is a second example of a call interface provided in an embodiment of this application;
[0018] Figure 5 is a third example of a call interface provided in an embodiment of this application;
[0019] Figure 6 is a fourth example of a call interface provided in an embodiment of this application;
[0020] Figure 7 is a schematic diagram of the structure of a communication device provided in an embodiment of this application;
[0021] Figure 8 is one of the hardware structure diagrams of an electronic device provided in an embodiment of this application;
[0022] Figure 9 is a second schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0023] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0024] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0025] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."
[0026] The following description, in conjunction with the accompanying drawings, details the call method, apparatus, electronic device, and storage medium provided in the embodiments of this application through specific examples and application scenarios.
[0027] The call method, apparatus, electronic device, and storage medium provided in this application can be applied to telephone call scenarios, video call scenarios, voice call scenarios, voice message receiving scenarios, voice recording scenarios, etc.
[0028] In one scenario, for call scenarios, such as the aforementioned telephone call scenario, video call scenario, or voice call scenario,
[0029] When User A and User B are talking, if User A outputs Chinese speech through Electronic Device 1, and Electronic Device 1 receives English speech from User B outputting through Electronic Device 2, then Electronic Device 1 can translate the English speech output by Electronic Device 2 into Chinese speech. User A can then understand User B's conversation through the Chinese speech translation. Furthermore, Electronic Device 1 can also translate User A's Chinese speech into English speech and send it to Electronic Device 2, so that User B can understand User A's conversation through the English translation.
[0030] In another scenario, for voice message reception, after an application (such as a chat application) on an electronic device receives a voice message, a message identifier for the voice message can be displayed in the chat interface. If the voice message contains English speech, the user can input the message identifier to allow the electronic device to display a translation control. The user can then input the translation control to allow the electronic device to translate the speech contained in the voice message and output the corresponding Chinese translation.
[0031] In another scenario, for voice recording, if the user records English voice using audio recording software on the electronic device, and wants to change the language of the recorded voice, the user can long-press the voice icon corresponding to the English voice to input it, so that the electronic device can display at least one language control. If the user selects the Chinese language control from at least one language control, the electronic device can translate the English voice into Chinese voice.
[0032] The call method, apparatus, electronic device, and storage medium provided in this application can translate the second speech and output the translated speech, enabling the first party to clearly understand the content of the second party's call directly based on the translated speech. Furthermore, by sending the translated speech of the first speech to the second party, the second party can clearly understand the content of the first party's call, thus avoiding communication barriers between the two parties and improving the communication efficiency of both parties.
[0033] The execution subject of the call method provided in this application embodiment can be a call device, which can be an electronic device or a functional module in an electronic device. The following uses an electronic device as an example to illustrate the technical solution provided in this application embodiment.
[0034] This application provides a call method. Figure 1 shows a flowchart of a call method provided by this application, applied to an electronic device corresponding to a first party in the call. As shown in Figure 1, the call method provided by this application may include the following step 201.
[0035] Step 201: In the case of a conversation between the first and second parties, perform at least one of the following:
[0036] Upon receiving the first voice message from the first party in the call, the first voice message and its translation, or the translation of the first voice message, are sent to the second party in the call; upon receiving the second voice message from the second party in the call, the second voice message and its translation, or the translation of the second voice message, are played.
[0037] Optionally, in this embodiment of the application, when the first party and the second party are talking, and the electronic device receives the first voice message from the first party, the electronic device can send the first voice message and a translated version of the first voice message, or a translated version of the first voice message, to the second party.
[0038] Optionally, in this embodiment of the application, when the first party and the second party are talking, and the electronic device receives the second voice from the second party, it can send the translated voice of the first voice to the second party; when the second voice from the second party is received, it can play the second voice and the translated voice of the second voice, or the translated voice of the second voice.
[0039] Optionally, in this embodiment of the application, when the first party and the second party are talking, and the electronic device receives the first voice message from the first party, it can send the first voice message and a translated version of the first voice message, or the translated version of the first voice message, to the second party; and when the second party receives the second voice message, it can play the second voice message and a translated version of the second voice message, or the translated version of the second voice message. The specific implementation can be determined according to actual usage requirements, and this embodiment of the application does not impose any limitations.
[0040] Optionally, in the embodiments of this application, the above-mentioned call process can be described in detail in the following embodiments, and will not be repeated here to avoid repetition.
[0041] Optionally, in the embodiments of this application, step 201 above can be specifically implemented by step 201a below.
[0042] Optionally, upon receiving a translated version of the second voice from the second party in the call, the translated version of the second voice or the translated version of the second voice may be played.
[0043] Optionally, if the first party in the call is the voice sender and the second party in the call is the voice receiver, the electronic device corresponding to the first party in the call can translate the first party's voice and send the first party's voice and the translated voice, or the translated voice, to the electronic device corresponding to the second party in the call.
[0044] Optionally, if the first party in the call is the voice sender and the second party in the call is the voice receiver, the electronic device corresponding to the first party in the call can send the first voice of the first party in the call to the electronic device corresponding to the second party in the call. The electronic device corresponding to the second party in the call can translate the first voice of the first party in the call and play the first voice of the first party in the call and the translated voice of the first voice, or the translated voice of the first voice.
[0045] Optionally, if the first party in the call is the voice receiver and the second party in the call is the voice sender, the electronic device corresponding to the first party in the call can translate the second party's second voice and play the translated voice of the second voice or the translated voice of the second voice.
[0046] Optionally, when the first party in the call is the voice receiver, the second party in the call is the voice sender. The electronic device corresponding to the second party in the call can translate the second party's second voice and send the second voice and the translated voice, or the translated voice, to the electronic device corresponding to the first party in the call. The electronic device corresponding to the first party in the call then plays the second voice and the translated voice, or the translated voice.
[0047] Step 201a: When the first party and the second party are talking, and the electronic device receives the first voice message from the first party, the first voice message from the first party and a translated voice message of the first voice message, or a translated voice message of the first voice message, are sent to the second party.
[0048] In this embodiment of the application, the first caller can be a user of an electronic device, and the second caller can be a user of another electronic device.
[0049] Optionally, in this embodiment of the application, the electronic device can establish a call connection with other electronic devices through the telephone number corresponding to other electronic devices, and then make a call with the second party.
[0050] Optionally, in this embodiment of the application, the electronic device can establish a call connection with the electronic device of the second party through a contact identifier in an instant messaging application, and then conduct a call with the second party.
[0051] For example, an electronic device runs an instant messaging application and then displays an instant messaging interface including at least one contact identifier. A first party can click on the first contact identifier to display voice call controls and video call controls. Then, the first party can click on the voice call control to establish a voice call connection with a second party's electronic device and then conduct a voice call with the second party. Alternatively, the first party can click on the video call control to establish a video call connection with a second party's second electronic device and then conduct a video call with the second party.
[0052] Optionally, in this embodiment, the contact identifier may include any of the following: Chinese characters, English characters, images, special symbols, or emoticons. The specific identifier can be determined according to actual usage needs, and this embodiment does not impose any limitations.
[0053] In this embodiment of the application, the first voice is an audio recording containing the voice of the first party in the call.
[0054] In this embodiment of the application, the first voice message may be received through the microphone of an electronic device.
[0055] Optionally, in this embodiment, the microphone can be any of the following: a condenser microphone, a piezoelectric microphone, a linear microphone, or a differential pressure microphone, etc. The specific type can be determined based on actual usage, and this embodiment does not impose any limitations.
[0056] Optionally, in this embodiment of the application, the above-mentioned call can be a voice call or a video call.
[0057] For example, when the first and second parties are in a voice call, the first voice message can be something like "Hello," "What are you doing?" or "Have you finished your work for today?" The specific message can be determined according to actual usage needs, and this application embodiment does not impose any limitations.
[0058] Optionally, in this embodiment of the application, the electronic device can send the first voice message to the electronic device of the second party in the call via a wireless network.
[0059] Optionally, in the embodiments of this application, the wireless network can be any of the following: a mobile network or a Wireless Fidelity (WiFi) network.
[0060] For example, the mobile network mentioned above can be any of the following: a 4G network or a 5G network.
[0061] Optionally, in the embodiments of this application, step 201 above can be specifically implemented by step 201b below.
[0062] Step 201b: When the first and second parties are talking and the electronic device receives the second voice message from the second party, play the second voice message and a translation of the second voice message, or a translation of the second voice message.
[0063] Optionally, in this embodiment of the application, when the first and second parties are talking, the electronic device can receive the second voice message through a telephone application; or, when the first and second parties are talking, the first electronic device can receive the second voice message through a chat application.
[0064] For example, in the case of a conversation between the first and second parties, the second voice message could be something like "Hello," "I'm working," or "I'll finish soon." The specific message can be determined based on actual usage needs, and this application embodiment does not impose any limitations.
[0065] Optionally, in this embodiment of the application, the electronic device can receive a second voice message via a wireless network.
[0066] In this embodiment of the application, the electronic device can output the second voice and the translated voice of the second voice, or the translated voice of the second voice, through a microphone.
[0067] Optionally, in this embodiment of the application, when displaying an incoming call interface, the incoming call interface may display a language translation control, which the user can click on to enable the electronic device to translate the first voice and the second voice.
[0068] For example, as shown in Figure 2, taking a mobile phone as an example, the mobile phone displays an incoming call interface 10. The incoming call interface 10 includes the incoming call number, the location of the phone number, and a language translation control 11. The user can click on the language translation control 11 to input, so that the mobile phone can translate the voice of the first caller and the voice of the second caller when the call is connected.
[0069] In the call method provided in this application embodiment, when a first party and a second party are talking, if a first voice message from the first party is received, the first voice message and its translation, or the translation of the first voice message, are sent to the second party; if a second voice message from the second party is received, the second voice message and its translation, or the translation of the second voice message, are played. In this solution, since the second voice message can be translated to output its translation, the first party can more clearly understand the semantics of the received voice message directly from the translation. Furthermore, sending the translation of the first voice message to the second party allows the second party to more clearly understand the semantics of the received voice message directly from the translation, avoiding communication barriers between the two parties and thus improving communication efficiency.
[0070] Optionally, in this embodiment of the application, the call method provided in this embodiment of the application further includes the following step 301.
[0071] Step 301: The electronic device determines the translated speech and the translated content of the first speech based on the language of the second party in the call.
[0072] In this embodiment of the application, the translated content can be translated text.
[0073] Optionally, in this embodiment of the application, the language of the second caller is determined based on at least one of the following: the Internet Protocol (IP) home location of the second caller, the second voice of the second caller, the contact settings information of the second caller, and the language selection control in the call interface between the first caller and the second caller.
[0074] In one example, the language of the second party in the call can be determined based on the second speech of the second party in the call: after receiving the second speech, the electronic device can determine the language of the second speech based on the second speech, and translate the first speech based on the language of the second speech to obtain the translated speech and the translated content of the first speech.
[0075] Optionally, in this embodiment, the language can be any of the following: Chinese, English, Russian, Spanish, or German, etc. The specific language can be determined according to actual usage requirements, and this embodiment does not impose any limitations.
[0076] In this embodiment of the application, the electronic device can input the second speech into a Hidden Markov Model (HMM) to determine the language of the second speech.
[0077] For example, an electronic device can input a second voice message into a Hidden Markov Model (HMM). The HMM segments the voice message to obtain at least one segment, then extracts Mel-frequency cepstral coefficients (MFCs) from each segment. These MFCs are then compared with MFCs stored in a database to determine the language corresponding to each MFC. It should be noted that one MFC corresponds to one language in the database.
[0078] In this embodiment of the application, the electronic device can input the first voice into a speech recognition algorithm, such as an Automatic Speech Recognition (ASR) algorithm, to obtain the text corresponding to the first voice, and then use a language translation algorithm to translate the text corresponding to the first voice according to the language of the second party in the conversation, so as to obtain the translated text of the first voice.
[0079] For example, an electronic device can input the first speech into an ASR algorithm. The ASR algorithm splits the first speech into audio frames, transforms the split audio frames into multi-dimensional vector information according to human ear characteristics, combines the multi-dimensional vector information to form phonemes, and finally combines the phonemes into words and strings them together to form sentences, thereby obtaining the text corresponding to the first speech. Then, through a language translation algorithm, the semantic information of the text corresponding to the first speech is obtained, and according to the language of the second speaker, the text representation of the semantic information in the language of the second speech is obtained, thereby obtaining the translated text of the first speech.
[0080] For example, if the language of the second party in the call is English, after receiving the first voice "Hello", the electronic device can first convert the first voice into a text representation, namely "Hello"; then, the first electronic device can translate the text "Hello" according to English to obtain the translated text "hello".
[0081] Optionally, in this embodiment of the application, the electronic device can translate the first speech using a first algorithm and the language of the second party in the call to obtain the translated speech corresponding to the first speech.
[0082] Optionally, in the embodiments of this application, the first algorithm described above can be any of the following: Artificial Intelligence (AI) algorithm, neural network algorithm, and large model algorithm, etc. The specific algorithm can be determined according to the actual application, and the embodiments of this application do not impose any limitations.
[0083] In another example, the language of the second party can be determined based on the language selection control in the call interface between the first and second parties: the call interface can display at least one language control, and the language corresponding to the first language control is determined as the language of the second party by the user's selection input of the first language control among the at least one language control.
[0084] In this embodiment of the application, each of the at least one language control is used to determine a language, that is, each of the at least one language control can correspond to a language.
[0085] Optionally, in this embodiment of the application, the electronic device may display at least one language control in a blank area of the call interface; or, during a call, the first electronic device may display at least one language control in the call interface via a pop-up window.
[0086] In this embodiment of the application, the above-mentioned selection input is used to select a first language control from at least one language control.
[0087] Optionally, in the embodiments of this application, the above-mentioned input selection includes, but is not limited to: user clicking on the first language control through a touch device such as a finger or stylus, or voice commands input by the user, or specific gestures input by the user, or other feasible inputs. The specific input can be determined according to actual usage needs, and the embodiments of this application do not limit it.
[0088] In some embodiments of this application, the aforementioned specific gesture can be any one of a single-click gesture, a swipe gesture, a drag gesture, a pressure-recognition gesture, a long-press gesture, an area-changing gesture, a double-press gesture, or a double-tap gesture.
[0089] In some embodiments of this application, the aforementioned click input can be a single click, a double click, or any number of clicks, and can also be a long press or a short press. For example, the aforementioned first input can be: a user's click input on a first language control.
[0090] For example, as shown in Figure 3, the mobile phone can display three language controls on the call interface 12, which respectively display Chinese, English, and Russian. The user can click on the English language control to input, so that the electronic device can determine the language corresponding to the English language control as the language of the second party in the call.
[0091] In another example, the language of the second caller can be determined based on the contact settings information of the second caller: the electronic device displays a contact interface, which includes at least one contact identifier; the electronic device receives a third input to a target contact identifier among the at least one contact identifier; in response to the second input, displays at least one language control corresponding to the target contact identifier; receives a fourth input from the user to a second language control among the at least one language controls; and in response to the fourth input, determines the language corresponding to the second language control as the language of the second caller.
[0092] It is understood that each of the above-mentioned contact identifiers is used to indicate a contact.
[0093] Optionally, in this embodiment of the application, the contact interface can be a contact interface in a phone application or a contact interface in an instant messaging application.
[0094] For example, an electronic device can run a phone application and display a contacts interface based on a user's click on the phone application icon.
[0095] As another example, an electronic device can run a phone application and display a contacts interface based on the user's click input on an instant messaging application icon.
[0096] In this embodiment of the application, the third input is used to select a target contact identifier from at least one contact identifier.
[0097] Optionally, in this application embodiment, the aforementioned third input includes, but is not limited to: the user clicking on the target contact identifier using a touch device such as a finger or stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs. The specific input can be determined according to actual usage needs, and this application embodiment does not limit it.
[0098] Optionally, in this embodiment of the application, the electronic device may display at least one language control corresponding to the target contact identifier in the contact interface; or, the electronic device may display at least one language control corresponding to the target contact identifier through a pop-up window.
[0099] In this embodiment of the application, the fourth input is used to select a second language control from at least one language control.
[0100] Optionally, in the embodiments of this application, the fourth input includes, but is not limited to: the user clicking on the second language control through a touch device such as a finger or stylus, or the user's voice command, or the user's specific gesture, or other feasible inputs. The specific input can be determined according to actual usage needs, and the embodiments of this application do not limit it.
[0101] In this embodiment of the application, the electronic device can pre-set the language of the second party in the call, so that during the call, the electronic device can directly translate the first speech according to the language of the second party in the call.
[0102] In another example, the language of the second party in the call can be determined based on the location of the second party's Internet Protocol (IP). For example, if the electronic device determines that the IP address is located in Beijing, it can determine that the language of the second party in the call is Chinese based on the rule that Beijing belongs to China and the common language of China is Chinese. Alternatively, if the electronic device determines that the IP address is located in Guangdong, the first electronic device can determine that the language of the second party in the call is Chinese based on the rule that Guangdong belongs to China and the common language of China is Chinese.
[0103] In this embodiment, the electronic device can send the translated speech of the first voice to the second party in the call, enabling the second party to understand the semantics of the received speech more clearly based on the translated speech, avoiding communication barriers between the two parties and thus improving the communication efficiency of both parties.
[0104] Optionally, in this embodiment of the application, the call method provided in this embodiment of the application further includes the following step 401.
[0105] Step 401: The electronic device determines the translated speech and the translated content of the second speech based on the language of the first party in the call.
[0106] Optionally, in this embodiment of the application, the language of the first caller can be determined based on at least one of the following: the location of the Subscriber Identity Module (SIM) of the electronic device, the first voice of the first caller, and the language selection control in the call interface between the first caller and the second caller.
[0107] It should be noted that the specific implementation process described above can be found in the above embodiments, and will not be repeated here to avoid repetition.
[0108] Optionally, in this embodiment of the application, when the first caller and the second caller are in a call, at least one of the following is displayed in the call interface between the first caller and the second caller: the voice content of the first voice, the translated content of the first voice, the voice content of the second voice, and the translated content of the second voice.
[0109] For example, as shown in Figure 4, when the call interface 12 is displayed, the mobile phone can display the original text corresponding to the first voice, "Is there anything else I can help you with?" and the translated text corresponding to the first voice, "Is there anything else I can help you with?", in the call interface 12. At this time, the second voice is "No more for now, thank you very much". Then the mobile phone can display the Chinese "No more for now, thank you very much" in the call interface 12, and display the translated text "No more for now, thank you very much" below the Chinese "No more for now, thank you very much".
[0110] Optionally, in this embodiment of the application, during a typical call, the electronic device is held close to the ear, which may make it inconvenient to view the translated content. If the electronic device detects that the device is close to the ear, it may not display the translated content; if it detects that the device is removed from the ear, it will display the translated content, thus saving energy.
[0111] In this embodiment, since the second speech can be translated to output the translated speech, the first party in the call can more clearly understand the semantics of the received speech by directly translating the second speech, thus avoiding communication barriers between the two parties and improving the communication efficiency of the two parties in the call.
[0112] Optionally, in this embodiment of the application, the number of the second parties to the call is at least two.
[0113] For example, before "the electronic device sends the first voice of the first caller and the translated voice of the first voice, or the translated voice of the first voice, to the second caller" in step 201 above, the call method provided in this application embodiment further includes the following step 501.
[0114] Step 501: The electronic device determines the translated speech of the first speech corresponding to each second speaker based on the language of each second speaker.
[0115] Optionally, in this embodiment of the application, the language of each of the second parties can be determined based on at least one of the following: the Internet Protocol (IP) home location of each of the second parties, the second voice of each of the second parties, the contact settings information of each of the second parties, and the language selection control in the call interface between the first party and each of the second parties.
[0116] In this embodiment of the application, the electronic device can translate the first speech according to the language of each second caller to obtain the translated speech corresponding to each second caller, and send the translated speech corresponding to each second caller to each second caller.
[0117] It should be noted that the specific implementation process can be found in the above embodiments, and will not be repeated here to avoid repetition.
[0118] In this embodiment, the electronic device can send the translated speech of the first speech to each of the second parties in the call, enabling each of the second parties to understand the semantics of the received speech more clearly based on the translated speech, avoiding communication barriers in multi-party calls, and thus improving the communication efficiency of multi-party calls.
[0119] Optionally, in this embodiment of the application, the number of the second parties to the call is at least two.
[0120] For example, before "the electronic device plays the second voice and the translated voice of the second voice, or the translated voice of the second voice" in step 201 above, the call method provided in this application embodiment further includes the following step 601.
[0121] Step 601: The electronic device determines the translated speech of the second speech corresponding to each of the second speakers based on the language of the first speaker.
[0122] In this embodiment of the application, the electronic device can translate the second speech sent by each of the second parties according to the language of the first party in the call, so as to obtain the translated speech corresponding to the second speech of each of the second parties in the call.
[0123] It should be noted that the specific implementation process can be found in the above embodiments, and will not be repeated here to avoid repetition.
[0124] In this embodiment, since each second speech can be translated to output translated speech, the first party in the call can more clearly understand the semantics of the received speech by directly translating the second speech, thus avoiding communication barriers in multi-party calls and improving the communication efficiency of multi-party calls.
[0125] Optionally, in this embodiment of the application, before "the electronic device sends the first voice of the first caller and the translated voice of the first voice, or the translated voice of the first voice, to the second caller" in step 201 above, the call method provided in this embodiment of the application further includes the following step 701.
[0126] Step 701: The electronic device generates a translated speech of the first speech based on the voiceprint characteristics of the first caller.
[0127] In this embodiment of the application, the aforementioned voiceprint feature can be a sound wave spectrum carrying speech information, and the voiceprint feature can characterize the user's voice.
[0128] Optionally, in the embodiments of this application, the above-mentioned speech information includes at least one of the following: timbre, pitch, intensity, and duration.
[0129] Optionally, in this embodiment of the application, the call interface includes a simultaneous interpretation control. The first caller can input into the simultaneous interpretation control so that the first electronic device can output the translated speech of the first speech based on the voiceprint features corresponding to the first speech.
[0130] For example, an electronic device can add voiceprint features to AI voice information to obtain a new voice, and then use the new voice as the voice of the translated voice of the first voice, and output the translated voice through the voice.
[0131] Optionally, in this embodiment of the application, when the first party in the call does not input the simultaneous interpretation control, the electronic device can output the translated speech of the first speech through AI voice.
[0132] In this embodiment of the application, the electronic device can obtain the voiceprint features corresponding to the first voice through a second algorithm.
[0133] Optionally, in the embodiments of this application, the second algorithm described above can be any of the following: an AI algorithm, a neural network algorithm, or a voiceprint extraction model.
[0134] For example, the first electronic device extracts features from multiple audio frames in the first speech using a voiceprint extraction model to obtain multiple target frame features corresponding to the multiple audio frames; then, based on the multiple target frame features, it obtains the covariance matrix, variance, and mean corresponding to the multiple audio frames; it performs dimensionality reduction processing on the covariance matrix corresponding to the multiple audio frames using a one-dimensional convolutional layer in the voiceprint extraction model to obtain the dimensionality-reduced result, and performs vectorization operation on the dimensionality-reduced result to obtain target one-dimensional preprocessed data; it performs concatenation operation on the variance, mean, and target one-dimensional preprocessed data corresponding to the multiple audio frames to obtain the target concatenation result; and it performs voiceprint recognition processing on the target concatenation result using the voiceprint extraction model to obtain the voiceprint features of the first speech.
[0135] In this embodiment, the electronic device outputs a translated version of the first speech based on the voiceprint features corresponding to the first speech, which makes the output translated speech sound more natural and thus improves the call experience.
[0136] Optionally, in this embodiment of the application, when the electronic device receives the second voice from the second party, it plays the translated voice of the second voice, and the original voice control in the call interface between the first party and the second party is in a closed state; or, when the electronic device receives the second voice from the second party, it plays the second voice and the translated voice of the second voice, and the original voice control in the call interface between the first party and the second party is in a closed state.
[0137] For example, referring to Figure 4 and as shown in Figure 5, the call interface 12 includes a first caller's original audio control 15 and a second caller's original audio control 16. The user can click on the first caller's original audio control 15 to stop the phone from playing the first caller's original audio; and the user can click on the second caller's original audio control 16 to stop the phone from playing the second caller's original audio.
[0138] Optionally, in this embodiment of the application, the call method provided in this embodiment of the application further includes the following step 801.
[0139] Step 801: The electronic device displays commonly used phrases corresponding to the call scenario in the call interface between the first and second callers.
[0140] In this embodiment of the application, the electronic device can determine the call scenario based on the first voice and the second voice.
[0141] For example, if the second voice message is "Good evening, sir, what cuisine would you like to order?" and the first voice message is "Order Shandong cuisine", the mobile phone can determine the keywords "order, cuisine" based on the second and first voice messages, and thus determine that the call scenario is a food ordering scenario.
[0142] For example, if the second voice message is "What time do I need to check in?" and the first voice message is "3 p.m.", the mobile phone determines the keyword "check-in" based on the second and first voice messages, thus determining that the call scenario is a hotel booking scenario.
[0143] In this embodiment of the application, there is a correlation between the above-mentioned call scenario and commonly used words. That is to say, after the first electronic device determines the call scenario, the first electronic device can directly determine the commonly used words based on the call scenario.
[0144] It is understandable that the above-mentioned commonly used words are words frequently used by the first party in a certain context. There may be one or more of these commonly used words.
[0145] For example, taking a phone call scenario as an example of ordering food, the aforementioned common phrases could be "less salt, less oil".
[0146] For example, as shown in Figure 6, when the first and second parties are in a call, the mobile phone can display the commonly used words "don't eat spicy food" and "want a private room" on the call interface based on the voice content of both parties.
[0147] For example, taking a hotel booking scenario as an example, the aforementioned commonly used term could be "with window".
[0148] Optionally, in this embodiment of the application, when the electronic device detects that the call scenario includes information such as time, location, people and events, the electronic device can automatically create a pending matter and automatically create an alarm clock based on the aforementioned time to remind the first caller to handle the pending matter.
[0149] Optionally, in this embodiment of the application, the call interface may further include a text input box, in which the first party can input text, so that the electronic device can send the text to the electronic device of the second party.
[0150] Optionally, in this embodiment of the application, the user can long-press to input the above-mentioned commonly used words so that the electronic device can display a commonly used word editing interface, where the user can add, delete, or edit commonly used words.
[0151] In this embodiment, issues such as noisy environments or different speaking styles may occur, leading to translation errors or speech errors. To avoid such problems, the call interface provides commonly used phrases for the first caller, allowing them to directly use these phrases so that the second caller can clearly understand the first caller's meaning.
[0152] Optionally, in this embodiment of the application, the call method provided in this embodiment of the application further includes the following steps 901 and 902.
[0153] Step 901: The electronic device receives the first input of the translated content of the first speech.
[0154] In this embodiment of the application, the first input is the input for editing the translation content of the first speech.
[0155] Optionally, in the embodiments of this application, the first input includes, but is not limited to: text input of the translation content of the first speech by the user through a touch device such as a finger or stylus, or voice command input by the user, or specific gesture input by the user, or other feasible inputs. The specific input can be determined according to actual usage needs, and the embodiments of this application do not limit it.
[0156] Step 902: The electronic device responds to the first input by updating the translated content of the first speech and the translated speech of the first speech.
[0157] In this embodiment of the application, the electronic device can update the translation content of the first speech by the user's editing of the translation content of the first speech, and then update the translated speech of the first speech with the updated translation content of the first speech.
[0158] In this embodiment, the electronic device can send the updated translation of the first speech and the translated speech of the first speech to the second party in the call.
[0159] In this embodiment, the electronic device can correct the translation of the first speech based on user input and send the corrected translation of the first speech to the second party in the call, thereby improving the accuracy of the electronic device in translating the first speech.
[0160] Optionally, in the embodiments of this application, the call method provided in the embodiments of this application further includes the following steps 1001 and 1002.
[0161] Step 1001: The electronic device receives a second input of the translated content of the second speech.
[0162] In this embodiment of the application, the second input is the input for selecting the translation content of the first speech.
[0163] Optionally, in this embodiment, the second input includes, but is not limited to: the user clicking on the translated content of the second speech using a touch device such as a finger or stylus, or a voice command input by the user, or a specific gesture input by the user, or other feasible inputs. The specific input can be determined according to actual usage needs, and this embodiment does not limit it.
[0164] Step 1002: The electronic device responds to the second input by sending a repeat query request to the second party in the call.
[0165] In this embodiment of the application, after the electronic device sends a repeat query request to the electronic device of the second party, the second party can resend the second voice message.
[0166] In this embodiment of the application, the electronic device sends a repeat query request to the second party in the call, so that the second party in the call can resend the second voice message, so that the first party in the call can clearly understand the meaning of the second voice message.
[0167] It should be noted that the call method provided in this application embodiment can be executed by a call device. This application embodiment uses a call device executing the call method as an example to illustrate the call device provided in this application embodiment.
[0168] Figure 7 shows a possible structural schematic diagram of the communication device involved in an embodiment of this application. As shown in Figure 7, the communication device 70 may include: a sending module 71 and a playback module 72.
[0169] The sending module 71 is used to send the first voice message and its translation, or the translation of the first voice message, to the second voice message when the first and second voice messages are being received. The playback module 72 is used to play the second voice message and its translation, or the translation of the second voice message, when the second voice message is received from the second voice message.
[0170] In one possible implementation, the playback module 72 is further configured to play the second voice and the translated voice of the second voice, or the translated voice of the second voice, upon receiving the translated voice of the second voice from the second caller.
[0171] In one possible implementation, when the first and second parties are in a conversation, at least one of the following is displayed on the call interface between the first and second parties: the voice content of the first voice, the translated content of the first voice, the voice content of the second voice, and the translated content of the second voice.
[0172] In one possible implementation, the playback device 70 provided in this application embodiment further includes a determining module. The determining module is configured to determine the translated speech and the translated content of the first speech based on the language of the second party in the call.
[0173] In one possible implementation, the playback device 70 provided in this application embodiment further includes a determining module. The determining module is configured to determine the translated speech of the second speech and the translated content of the second speech based on the language of the first party in the call.
[0174] In one possible implementation, the number of the second parties to the call is at least two; the playback device 70 provided in this application embodiment further includes: a determining module. The determining module is configured to, before sending the first voice and its translated voice, or the translated voice, of the first party to the second party to the second party, determine the translated voice corresponding to the first voice for each of the second parties, based on the language of each second party.
[0175] In one possible implementation, the number of the second parties in the call is at least two; the playback device 70 provided in this application embodiment further includes: a determining module. The determining module is used to determine, based on the language of the first party, the translated speech of the second speech corresponding to each of the second parties before playing the second speech or the translated speech of the second speech.
[0176] In one possible implementation, the playback device 70 provided in this application embodiment further includes a generation module. The generation module is configured to generate a translated voice of the first speech based on the voiceprint characteristics of the first party before sending the first speech and its translated form, or the translated voice, of the first party to the second party.
[0177] In one possible implementation, the language of the second party is determined based on at least one of the following: the IP address of the second party, the second voice of the second party, the contact settings of the second party, and the language selection control in the call interface between the first party and the second party.
[0178] In one possible implementation, upon receiving a second voice message from a second party, a translated version of the second voice message is played, and the original audio control in the call interface between the first and second parties is turned off; upon receiving a second voice message from a second party, a translated version of the second voice message is played, and the original audio control in the call interface between the first and second parties is turned on.
[0179] In one possible implementation, the playback device 70 provided in this application embodiment further includes a display module. The display module is used to display commonly used phrases corresponding to the call scenario in the call interface between the first and second parties.
[0180] In one possible implementation, the playback device 70 provided in this application embodiment further includes a receiving module and an updating module. The receiving module is configured to receive a first input of the translated content of the first speech. The updating module is configured to update the translated content of the first speech and the translated speech of the first speech in response to the first input.
[0181] In one possible implementation, the playback device 70 provided in this application embodiment further includes: a receiving module. The receiving module is configured to receive a second input of translated content of the second speech. The sending module is further configured to send a repeat query request to the second party in response to the second input.
[0182] This application provides a communication device that can translate a second voice message and output the translated voice message, enabling the first party in the call to clearly understand the content of the second party's call based on the translated voice message. Furthermore, by sending the translated voice message of the first voice message to the second party, the second party can clearly understand the content of the first party's call, thus avoiding communication barriers between the two parties and improving the communication efficiency of both parties.
[0183] The calling device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, a mobile electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.
[0184] The communication device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system used.
[0185] The communication device provided in this application embodiment can implement the various processes implemented in the above embodiments. To avoid repetition, it will not be described again here.
[0186] Optionally, as shown in FIG8, this application embodiment also provides an electronic device 90, including a processor 91 and a memory 92. The memory 92 stores a program or instructions that can run on the processor 91. When the program or instructions are executed by the processor 91, they implement the various steps of the above-described information output method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0187] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.
[0188] Figure 9 is a schematic diagram of the hardware structure of an electronic device that implements an embodiment of this application.
[0189] The electronic device 100 includes, but is not limited to, components such as: radio frequency unit 101, network module 102, audio output unit 103, input unit 104, sensor 105, display unit 106, user input unit 107, interface unit 108, memory 109, and processor 110.
[0190] Those skilled in the art will understand that the electronic device 100 may also include a power supply (such as a battery) for powering various components. The power supply can be logically connected to the processor 110 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The electronic device structure shown in Figure 9 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0191] The processor 110 is configured to perform at least one of the following when a first party and a second party are in a conversation: upon receiving a first voice message from the first party, sending the first voice message and a translated version of the first voice message, or a translated version of the first voice message, to the second party; upon receiving a second voice message from the second party, playing the second voice message and a translated version of the second voice message, or a translated version of the second voice message.
[0192] Optionally, in this embodiment of the application, the processor 110 is further configured to play the second voice and the translated voice of the second voice, or the translated voice of the second voice, when receiving the translated voice of the second voice from the second party in the call.
[0193] Optionally, in this embodiment of the application, the processor 110 is further configured to determine the translated speech and the translated content of the first speech based on the language of the second party in the call.
[0194] Optionally, in this embodiment of the application, the processor 110 is further configured to determine the translated speech of the second speech and the translated content of the second speech based on the language of the first party in the call.
[0195] Optionally, in this embodiment of the application, the number of the second parties to the call is at least two; the processor 110 is further configured to determine the translated voice of the first voice corresponding to each of the second parties to the call before sending the first voice and the translated voice of the first voice, or the translated voice of the first voice, to the second parties to the call.
[0196] Optionally, in this embodiment of the application, the number of the second parties to the call is at least two; the processor 110 is further configured to play the second speech and the translated speech of the second speech, or, before the translated speech of the second speech, determine the translated speech of the second speech corresponding to each of the second parties to the call based on the language of the first party to the call.
[0197] Optionally, in this embodiment of the application, the processor 110 is further configured to generate a translated voice of the first voice based on the voiceprint characteristics of the first party before sending the first voice and the translated voice of the first voice, or the translated voice of the first voice, of the first party to the second party.
[0198] Optionally, in this embodiment of the application, the display unit 106 is used to display commonly used phrases corresponding to the call scenario in the call interface between the first caller and the second caller.
[0199] Optionally, in this embodiment, the user input unit 107 is further configured to receive a first input of the translated content of the first speech. The processor 110 is further configured to update the translated content of the first speech and the translated speech in response to the first input.
[0200] Optionally, in this embodiment, the user input unit 107 is further configured to receive a second input of the translated content of the second speech. The processor 110 is further configured to send a repeat query request to the second party in response to the second input.
[0201] This application provides an electronic device that can translate a second voice message and output the translated voice message, enabling the first party in the call to clearly understand the content of the second party's call directly from the translated voice message; moreover, by sending the translated voice message of the first voice message to the second party, the second party can clearly understand the content of the first party's call, thus avoiding communication barriers between the two parties and improving the communication efficiency of both parties.
[0202] The electronic device provided in this application embodiment can implement the various processes implemented in the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0203] For details on the beneficial effects of the various implementation methods in this embodiment, please refer to the beneficial effects of the corresponding implementation methods in the above method embodiments. To avoid repetition, these will not be repeated here.
[0204] It should be understood that, in this embodiment, the input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 may include a touch detection device and a touch controller. Other input devices 1072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0205] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 109 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 109 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.
[0206] Processor 110 may include one or more processing units; optionally, processor 110 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 110.
[0207] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0208] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0209] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0210] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0211] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the information output method embodiments described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0212] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0213] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0214] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A call method, applied to an electronic device corresponding to a first party, the method comprising: In the event of a conversation between the first party and the second party, perform at least one of the following: Upon receiving the first voice message from the first party in the call, the first voice message from the first party in the call and a translated version of the first voice message, or a translated version of the first voice message, are sent to the second party in the call. Upon receiving the second voice message from the second party in the call, play the second voice message and its translation, or the translation of the second voice message.
2. The method according to claim 1, wherein, The method further includes: Upon receiving a translated version of the second voice from the second party in the call, play the second voice and its translated version, or the translated version of the second voice.
3. The method according to claim 1, wherein, When the first party and the second party are in a conversation, at least one of the following shall be displayed in the call interface between the first party and the second party: The audio content of the first audio, the translated content of the first audio, the audio content of the second audio, and the translated content of the second audio.
4. The method according to claim 3, wherein, The method further includes: Based on the language of the second party in the call, determine the translated speech of the first speech and the translated content of the first speech; Based on the language of the first party in the call, determine the translated speech of the second voice and the translated content of the second voice.
5. The method according to claim 1, wherein, The number of the second party in the call is at least two; Before sending the first voice message and its translated version, or the translated version of the first voice message, to the second party in the call, the method further includes: Based on the language of each of the second parties in the call, the translated speech of the first speech corresponding to each of the second parties in the call is determined.
6. The method according to claim 1, wherein, The number of the second party in the call is at least two; Before playing the second speech and its translated speech, or the translated speech of the second speech, the method further includes: Based on the language of the first party in the call, the translated speech of the second speech corresponding to each of the second parties in the call is determined.
7. The method according to claim 1, wherein, Before sending the first voice message and its translated version, or the translated version of the first voice message, to the second party in the call, the method further includes: The translated speech of the first speech is generated based on the voiceprint characteristics of the first caller.
8. The method according to claim 4, wherein, The language of the second party in the call is determined based on at least one of the following: the IP address of the second party in the call, the second voice of the second party in the call, the contact settings information of the second party in the call, and the language selection control in the call interface between the first party in the call and the second party in the call.
9. The method according to claim 1, wherein, Upon receiving a second voice message from the second party, the translated voice message is played, and the original voice control in the call interface of the first party and the second party is turned off; upon receiving a second voice message from the second party, the second voice message and its translated voice message are played, and the original voice control in the call interface of the first party and the second party is turned on.
10. The method according to claim 3, wherein, The method further includes: The call interface between the first and second parties displays commonly used phrases corresponding to the call scenario.
11. The method according to claim 3, wherein, The method further includes: Receive a first input of the translated content of the first speech; In response to the first input, update the translated content of the first speech and the translated speech of the first speech.
12. The method according to claim 3, wherein, The method further includes: Receive a second input containing the translated content of the second speech; In response to the second input, a repeat query request is sent to the second party in the call.
13. A communication device, applied to an electronic device corresponding to a first party in a communication, the communication device comprising: Sending module and playback module; The sending module is used to send the first voice and its translation, or the translation of the first voice, to the second voice when the first party and the second party are talking. The playback module is used to play the second voice and its translation, or the translation of the second voice, when it receives the second voice from the second party in the call.
14. An electronic device comprising a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the call method as claimed in any one of claims 1 to 12.
15. A readable storage medium storing a program or instructions that, when executed by a processor, implement the steps of the call method as described in any one of claims 1 to 12.
16. A computer program product, said program product being executed by at least one processor to implement the call method as described in any one of claims 1 to 12.
17. A user equipment (UE) comprising a UE configured to perform a call method as claimed in any one of claims 1 to 12.
Citation Information
Patent Citations
Translation method and terminal
CN109286725A
Translation method based on voice communication and electronic equipment
CN109582976A
Speech translation method, communication terminal and computer readable storage medium
CN110472254A
Call voice translation method and device and earphone equipment
CN113286217A
Call method and device, electronic equipment and readable storage medium
CN119363881A