Audio translation method, electronic device, storage medium and product

By automatically adjusting the presentation of translation results based on the audio source, the problem of cumbersome operation in existing translation applications is solved, achieving efficient presentation of translation results and improving user experience.

CN119274537BActive Publication Date: 2026-08-04BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2024-11-14
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing translation applications are cumbersome to operate in real-time translation scenarios, resulting in low translation efficiency and an inability to automatically adjust the presentation of translation results based on the audio source, thus affecting user experience.

Method used

The system automatically determines how the translation results are presented based on the source of the audio, improving translation efficiency and meeting the needs of different translation scenarios by playing or displaying the translated audio/text by default.

Benefits of technology

It improves translation efficiency and enhances the user experience, especially in face-to-face conversations and in-device audio scenarios, providing translation results quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274537B_ABST
    Figure CN119274537B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio translation method, an electronic device, a storage medium and a product, and relates to the technical field of computers. The audio translation method comprises: translating original audio obtained by a user device; determining whether the original audio is from outside or inside the user device; in response to the original audio being from outside the user device, defaulting to playing translated audio of the original audio; and in response to the original audio being from inside the user device, defaulting to not playing the translated audio of the original audio and displaying translated text of the original audio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, and in particular to an audio translation method, electronic device, storage medium, and product. Background Technology

[0002] With the development of machine learning technology, users can utilize various translation tools to translate text, documents, and speech. For example, an application can capture a user's first language speech and translate it into a second language. The translation result is then played back to the user or displayed on the user's device screen. Summary of the Invention

[0003] According to some embodiments of this disclosure, an audio translation method is provided, comprising: translating original audio acquired by a user device; determining whether the original audio originates from outside or inside the user device; in response to the original audio originating from outside the user device, playing a translated audio of the original audio by default; and in response to the original audio originating from inside the user device, not playing the translated audio of the original audio by default, and displaying the translated text of the original audio.

[0004] According to some embodiments of the present disclosure, an electronic device is provided, including: a memory; and a processor coupled to the memory, the processor being configured to execute an audio translation method of any embodiment of the present disclosure based on instructions stored in the memory.

[0005] According to some embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, performs the audio translation method of any embodiment described in the present disclosure.

[0006] According to some embodiments of this disclosure, a computer program product is provided that, when the computer program product is run on a computer, causes the computer to implement the audio translation method of any embodiment of this disclosure.

[0007] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0008] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure.

[0009] Figure 1 A flowchart illustrating an audio translation method according to some embodiments of the present disclosure is shown.

[0010] Figure 2(a) shows a schematic diagram of the interactive interface of an intelligent agent with translation capabilities.

[0011] Figure 2 (b) shows a schematic diagram of the interface for communicating with an agent.

[0012] Figure 3 (a) to 3(c) show schematic diagrams of the interactive interface of the translation function according to some embodiments of the present disclosure.

[0013] Figure 4 (a) and 4(b) show schematic diagrams of the user interface during the background operation of the agent.

[0014] Figure 5 A schematic diagram of an interface showing a translated text display method according to some embodiments of the present disclosure is shown.

[0015] Figure 6 A schematic diagram of the structure of an audio translation apparatus according to some embodiments of the present disclosure is shown.

[0016] Figure 7 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.

[0017] Figure 8 Block diagrams of electronic devices according to other embodiments of the present disclosure are shown.

[0018] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation

[0019] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.

[0020] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.

[0021] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".

[0022] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.

[0023] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0024] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0025] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0026] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0027] Analysis revealed that users may have various translation needs in real-time translation scenarios. For example, when users are having face-to-face conversations, they may want the app to help translate the dialogue; or, when listening to or watching foreign language content on their devices, they may want to receive real-time translation results. Some translation apps have limited functionality. Others, while offering more features, require users to manually configure settings or switch between functions. Because real-time translation requires a rapid start-up, cumbersome operations can slow down the translation process, resulting in low translation efficiency.

[0028] This disclosure provides an audio translation method, electronic device, storage medium, and product. Embodiments of this disclosure automatically determine the presentation method of the translation result based on the source of the original audio, thereby improving translation efficiency and enhancing user experience.

[0029] Figure 1 A flowchart illustrating an audio translation method according to some embodiments of this disclosure is shown. Figure 1 As shown, the audio translation method of this embodiment includes steps S102 to S108.

[0030] In step S102, the raw audio acquired by the user equipment is translated.

[0031] User devices can listen to acquired audio as raw audio, provided the user has authorized it. For example, they can listen to audio after the user enables real-time translation.

[0032] After obtaining the original audio, translation functions can be used for translation. For example, the content in the audio can be identified to obtain the original text in the original language, and then the original text can be translated into text in another language (i.e., a language different from the original language) (which can be called translated text). During the translation process, a translation model can be used, and the specific translation process will not be described in detail here. When determining the language used in the translated text, it can be based on the default settings or on the original language.

[0033] The translation process can be performed locally on the user's device. Alternatively, the original audio or text can be sent to other devices (such as servers) for processing, and the user's device can then receive the translation results returned by those devices, which may be translated text or audio.

[0034] The user device can monitor the acquired audio stream and perform real-time translation. That is, the user device can acquire each segment of raw audio from the stream, translate its content, and then acquire the next segment and continue translating. This achieves real-time translation.

[0035] In step S104, it is determined whether the original audio source is external or internal to the user device.

[0036] User devices acquire audio through both external and internal means. External audio refers to ambient sound captured via external audio input devices (such as an external microphone), such as the sound of face-to-face conversations or the sound of an in-person meeting. Internal audio refers to sound emitted by the user device itself, such as sound played by applications running on the device, including audio and video playback, and the sounds of other users in an online meeting. Internal audio can be captured through internal audio input devices (such as an internal microphone). The type of device or interface used to capture the audio determines whether the raw audio originates from the external or internal environment of the user device.

[0037] Sometimes, audio can be obtained from both external and internal sources of the user's device. In such cases, the user can be proactively asked to determine which source of the original audio should be translated. Alternatively, the user can pre-set the translation scenario or monitoring channel to determine in advance whether external or internal audio should be translated.

[0038] In step S106, in response to the original audio originating from outside the user device, the translated audio of the original audio is played by default.

[0039] Translation audio is audio obtained by converting translated text into speech; that is, translation audio includes the translation result of the content of the original audio.

[0040] Playing the translated audio by default means that the translated audio will be played to the user after the translation function is activated, without the user changing the playback settings. If the user does not wish to listen to the audio during use, they can manually turn off the playback function. Additionally, the translated text can also be displayed on the user's device for reading and reference, if needed.

[0041] If the original audio originates from outside the user's device, it indicates that the user is currently focused on sounds outside their device. Therefore, the translated audio can be played back, allowing the user to quickly obtain the translation result without looking at the device screen. For example, if a user doesn't understand language A, and they hear a foreigner speaking language A, the device can acquire the foreigner's audio, translate it, and play it back to the user. This allows the user to quickly understand the foreigner's speech and respond accordingly.

[0042] The translated audio can be played through the user's device or an external audio device connected to the user's device (such as headphones, speakers, wearable devices, etc.).

[0043] In step S108, in response to the fact that the original audio comes from inside the user device, the translated audio of the original audio is not played by default, and the translated text of the original audio is displayed.

[0044] The default setting to not play translation audio means that once the translation function is activated, the translation audio will not be played to the user unless the user changes the playback settings. Users can manually enable playback if they wish to listen to audio during use.

[0045] If the original audio originates from within the user's device, it means the user is currently focused on the sound emitted by the device itself. Since the user's device is playing audio and occupies at least one channel, the translation audio is not played by default; instead, the translation result is displayed. Thus, the user can obtain the translation result through the user's device screen.

[0046] For example, if a user does not understand language A, when the user's device listens to or watches content in language A, the device can display the translated text to the user instead of playing the translated audio. This maintains the user's immersion in the content being listened to or watched, and conveys the translation result in text form.

[0047] The above embodiments determine the default presentation mode of the translation results based on whether the audio collected by the user device originates from within or outside the user device. Specifically, for external raw audio, the translation result is played; for internal translation audio, the translation result is not played, but rather displayed. Therefore, the embodiments of this disclosure can use a translation result presentation mode adapted to the current translation scenario based on the source of the audio, allowing users to quickly obtain translation results suitable for the current scenario. Thus, translation efficiency can be improved, and the user experience enhanced.

[0048] As mentioned earlier, the acquisition of raw audio can be triggered by the translation function. For example, in response to the audio translation function being triggered, external and internal audio on the user's device can be monitored. There are various ways to trigger the audio translation function, such as launching a translation application or using a translation function within an application.

[0049] With the development of artificial intelligence technology, users can quickly obtain information by interacting with intelligent agents. An intelligent agent is an intelligent virtual object based on a machine learning model, which can automatically determine its response to user input. For example, a user can interact with an intelligent agent that has translation capabilities. After receiving input (such as text, speech, audio, etc.), the intelligent agent can use a machine learning model to obtain a translation result and return it to the user. Several exemplary embodiments of obtaining translation results by triggering an intelligent agent are described below.

[0050] In some embodiments, an intelligent agent is invoked via a specified external audio device to trigger the translation function. The invoked intelligent agent, provided by an application in the user device, is an intelligent agent with translation capabilities. The intelligent agent then runs in the foreground or background of the user device to acquire the original audio being listened to by the user device, translate it, and then return the translation result to the user.

[0051] External audio devices connected to the user device include headphones, speakers, wearable devices, etc., which can connect to the user device via wired or wireless means. The intelligent agent can operate in response to the activation of the external audio device; for example, after Bluetooth headphones are removed from their case and successfully connected to a mobile phone, the intelligent agent can automatically operate. Alternatively, the intelligent agent can be activated based on user commands after the external audio device has started operating. Once the audio device is activated, it or the user device listens for user commands, which can be voice commands, gesture commands, trigger commands for physical or virtual controls on the user device, etc. Taking headphones as an example, after the user wears the headphones and they are connected to the user device, the user can speak a specified voice command, such as "Hello, Intelligent Agent A." After recognizing the content of the voice, Intelligent Agent A and its translation function are activated.

[0052] When a user invokes the agent via a specified external audio device, if the feedback mode for the translation result is set to play the translated audio, the translated audio can be played via that external audio device by default. Of course, the translated audio can also be played through other means, and the playback channel and method can be modified based on the user's actions.

[0053] In some embodiments, the translation function of the agent is triggered by listening to external and internal audio on the user's device. For example, the user can trigger the agent's translation function by activating a control corresponding to the agent's translation function within the application. Figure 2 (a) A schematic diagram of the interactive interface of agent B with translation capabilities is shown. Figure 2As shown in (a), in interface 21, the user can engage in text-based conversations with agent B. Real-time translation can be activated in response to the user triggering control 211 in interface 21. Once activated, agent B acquires the audio being listened to by the user's device and translates it.

[0054] In some embodiments, when the agent is in a real-time call with the user and is in translation mode, it listens to both external and internal audio from the user's device. If the agent only has translation mode, it defaults to translation mode when a real-time call with the agent (e.g., a voice call or video call) is initiated. If the agent has multiple modes, the translation function is activated in response to selecting translation mode as the call mode during a call between the user and the agent. Figure 2 (b) A schematic diagram of the communication interface with agent C is shown. Figure 2 As shown in (b), in the call interface 22, in response to the call mode selection control "Scenario Selection" 221, the user can determine the current scenario as the real-time translation scenario, that is, select the real-time translation mode.

[0055] The above methods are merely illustrative examples of how to trigger translation functions. Those skilled in the art may choose other methods as needed.

[0056] In translation scenarios, sometimes multiple parties involved in a dialogue need to obtain the translation result, while other times only the current user wants the translation result in a language they do not understand. Embodiments of this disclosure can provide various translation methods to meet the needs of different scenarios.

[0057] Some embodiments of this disclosure can achieve mutual translation between multiple languages. In some embodiments, the language of the spoken content in the original audio is identified; the spoken content in the first language is translated into the second language; and the spoken content in the second language is translated into the first language. That is, the original audio contains multiple languages, and in this embodiment, each language can be translated. However, although this embodiment is described using the first and second languages ​​as examples, it is not limited to situations where only two languages ​​are used in the speaking scenario. For example, if a third language is present, the spoken content in the third language can be further translated into the first language.

[0058] Identifying the language of each piece of content can be achieved using speech recognition models or classification models, which will not be elaborated here.

[0059] In multilingual translation scenarios, the system can play translations of content in multiple languages. These multiple languages ​​can be all languages ​​involved in the original audio, or multiple languages ​​supported by the translation function. For example, for spoken content in the first language, the translated audio in the second language can be played; for spoken content in the second language, the translated audio in the first language can be played. Similarly, if the original audio contains more languages, the system can also play audio translations of other languages. For instance, if a third language is present, the spoken content in the third language can be further translated into the first language, and the translated audio in the first language can be played for the spoken content in the third language.

[0060] This method of playing translated content in multiple languages ​​can be applied to scenarios such as face-to-face translation. For example, if two people are chatting face-to-face in different languages ​​and do not understand each other's languages, the original audio can be translated into language 2 and played back, allowing the person speaking language 2 to understand what the other person is saying; conversely, the audio can be translated into language 1 and played back, allowing the person speaking language 1 to understand what the other person is saying.

[0061] In some embodiments, when the user device is not connected to headphones, i.e., when the translated audio is played through a speaker or loudspeaker, the translation results, which include content in multiple languages, can be played by default. This ensures that everyone involved in the original audio can hear the translation results, and each person can hear the translation in their own language. Of course, this setting can be modified according to the user's needs.

[0062] In multilingual translation scenarios, only the content in the user-selected language from the translation result can be played. In some embodiments, the user's selection of a target language is received, and the target language includes one or more; the audio corresponding to the target language content in the translation result is played. For example, if the user is using language 1 and is also participating in the conversation, and the user's speech is translated into language 2, while the speech of other users is translated into language 1, then the played translation result will only include the translations of the other users' speech. Thus, in some translation scenarios, when only the user needs to hear the translation result, and other speakers do not, only the audio in the target language selected by the user in the translation result can be played.

[0063] In some embodiments, when the user device is connected to headphones, i.e., when the translation audio is played through headphones, the user-selected translation audio, including the target language, can be played by default. Since it's generally unnecessary to play the audio to other users not wearing headphones when listening through headphones, playing only the translation results that the user can understand reduces the amount of audio playback and improves the user experience in this scenario. Of course, this setting can be modified according to the user's needs.

[0064] Some embodiments of this disclosure may translate only the content in the user-specified language; that is, translating all content is not necessary. In some embodiments, the user's selection of a target language may be received; the non-target language content in the original audio may be translated into the target language. For example, if the original audio includes content in language 1 and language 2, and the user selects language 1 as the target language, then only the content in language 1 of the original audio needs to be translated, while the content in language 2 does not need to be translated. Thus, the amount of content to be translated can be reduced according to the user's needs, saving computing resources and improving processing efficiency.

[0065] When the original audio originates from within the user's device, the translated text of the original audio is displayed. In some embodiments, the translated text of the original audio may also be displayed in response to the original audio originating from outside the user's device. This translated text may be displayed in real-time as the translation progresses, for example, when the user's device screen is on. Of course, when the original audio originates from outside the user's device, the user can obtain the translation result simply by listening to the audio; in this case, the user's device may be in a screen-off state, or the translation function may be running in the background. In this case, the translated text can be recorded in the background and displayed in response to the user's viewing of the translated text.

[0066] There are various ways to display translated text. Several examples are described below.

[0067] In some embodiments, in response to the agent running in the foreground, the translated text of the original audio is displayed in the agent's interactive interface. For example, for an agent with translation capabilities, such as a conversational agent, the translation result can be displayed in an interactive interface such as a message conversation with the agent or a real-time call interface.

[0068] In real-time translation scenarios, the user device continuously monitors multiple raw audio streams, i.e., it monitors the video stream. During translation, these raw audio segments can be processed in segments, and the playback result of each segment can be displayed or played one by one. Different segments can be divided according to duration, content volume, or speaker.

[0069] Taking duration as an example, the duration of the original audio or translated audio corresponding to each segment of translated text does not exceed a first threshold. Taking content volume as an example, the content volume of each segment of translated text is less than a threshold.

[0070] Taking the segmentation by speaker as an example, it can display one or more speakers from the original audio, as well as the translated text of the original audio corresponding to each speaker. Figure 3(a) A schematic diagram showing an interactive interface for translation functionality according to some embodiments of the present disclosure. Figure 3 As shown in (a), the translation results can be displayed in the interactive interface 30. For example, text comprising multiple paragraphs can be displayed, with paragraph 301 corresponding to speaker 1 and paragraph 302 corresponding to speaker 2. Speaker 1 uses Chinese, and paragraph 301 includes the original Chinese speech content 3011 and the English translation result 3012 of 3011. Speaker 2 uses English, and paragraph 302 includes the original English speech content 3021 and the Chinese translation result 3022 of 3021.

[0071] In some embodiments, in the agent's interactive interface, in response to the original audio originating from within the user's device and the user triggering a certain segment, the translated audio corresponding to that segment is played. Since the translated audio is not played by default when the original audio originates from within the user's device, the translated audio that the user wants to hear can be played by triggering a specific segment. For example, in response to triggering the playback identifier 3023 of segment 302, the translated audio of that segment can be played. During the playback of the translated audio, the identifier 3023 can switch to other styles to indicate that the associated segment is playing.

[0072] When the agent is running in the background, the translated text can be either not displayed or displayed.

[0073] In some embodiments, in response to the agent running in the background, a translation icon is displayed on the user device's interface when the original audio originates from outside the user device. When the original audio originates from outside the user device, playing the translated audio can satisfy the user's translation needs, and displaying the translated text is not necessary. Therefore, if the user needs to use the user device to view functions other than translation, only the translation icon may be displayed.

[0074] Figure 4 Figures (a) and (b) illustrate the user interface during the background operation of the agent. Figure 4 As shown in (a), interface 41 displays interfaces other than the translation function interface. This interface may include a translation indicator 411 to prompt the user that real-time translation is currently in progress. If needed, in response to triggering the translation indicator 411, the translation function interface may be displayed, for example, showing the translated text. Figure 4 (b) The interface 42 shown is the lock screen of the user's device. On the lock screen 42, a symbol 421 can be displayed to indicate to the user that real-time translation is currently in progress.

[0075] In some embodiments, when the agent is running in the background, a floating layer is displayed above the currently displayed interface; the translated text of the latest segment of the original audio is displayed in the floating layer. Thus, the translation result can be displayed to the user with minimal impact on the user's browsing of other application content. Especially when the currently displayed interface is the interface from which the audio source is located, the translated text displayed on the floating layer can function like subtitles, helping the user quickly understand the content currently being listened to or viewed.

[0076] Figure 5 A schematic diagram of an interface showing a translated text display method according to some embodiments of the present disclosure is shown. Figure 5 The interface shown is a podcast playback interface 50. The content played on this podcast interface is the original audio, which is technology news broadcast by host X (Host X). The language used in the podcast content is not the user's language. Therefore, the user can enable the real-time translation function. A floating layer 501 is displayed on top of the podcast interface, which includes the translated text of the podcast content so that the user can understand the podcast content in a timely manner. At the same time, the translation function can run in the background, and the user can continue to operate the podcast playback. Therefore, while the user efficiently obtains the translation results, it does not affect the efficiency of operating the currently focused application.

[0077] The above embodiments provide exemplary descriptions of the playback of translated audio and the display of translated text. The following describes several application scenarios to which the embodiments of this disclosure are applicable.

[0078] In an exemplary application scenario, a user device acquires audio from an external source, and headphones are not currently connected. This indicates a face-to-face translation scenario. The system identifies each language in the original audio, performs translations between multiple languages, and plays the translation results for each language. In this face-to-face translation scenario, all participants can understand each other's speech.

[0079] In an exemplary application scenario, a user device acquires audio from an external source, and headphones are currently connected. This indicates a situation where the user needs translated audio from their environment, such as attending an offline foreign language conference or lecture, or conversing with someone in a language the user doesn't understand. The system identifies language elements in the original audio that the user doesn't understand, and plays the corresponding translations to the user through the headphones, ensuring all translations are in the user's language.

[0080] For the two scenarios mentioned above, the translated text can be recorded in the background and displayed to the user when the user views it.

[0081] In one exemplary application scenario, the user device retrieves audio from its internal storage. This allows us to determine if the user needs to translate content from other foreign language applications on the device. In this case, the translation result can be displayed on the interface via a pop-up.

[0082] The above application scenarios provide some default configurations for providing translation results to users; however, those skilled in the art can adjust these configurations as needed. Furthermore, users can also make manual changes.

[0083] For example, in Figure 3 In the interactive interface 30 shown in (a), control 303 is used to set the translation mode. The currently used translation mode is Chinese-English translation, that is, translating the Chinese content in the original audio into English and the English content into Chinese. In response to triggering control 303, the following can be displayed: Figure 3 (b) contains a floating layer 3030, which includes multiple translation modes for the user to choose from. This allows the user to easily determine the target language for the translation.

[0084] For example, in Figure 3 In the interactive interface 30 shown in (a), control 304 is used to set the playback mode. The currently used playback mode is to not play the translated audio. In response to triggering control 304, the following can be displayed: Figure 3 (c) contains a floating layer 3040, which includes multiple playback modes for the user to choose from. This allows the user to easily determine which language translations will be played by default.

[0085] The methods of various embodiments of this disclosure have been described above. The apparatus for performing the methods of the above embodiments is described below.

[0086] Figure 6 A schematic diagram of the structure of an audio translation apparatus according to some embodiments of the present disclosure is shown. For example... Figure 6 As shown, the audio translation device 60 of this embodiment includes: a translation module 601 configured to translate the original audio acquired by the user device; a determination module 602 configured to determine whether the original audio comes from outside or inside the user device; and an interaction module 603 configured to, in response to the original audio coming from outside the user device, play the translated audio of the original audio by default; and in response to the original audio coming from inside the user device, not play the translated audio of the original audio by default, and display the translated text of the original audio.

[0087] In some embodiments, the audio translation device 60 further includes a monitoring module 604, configured to monitor external and internal audio of the user device in response to the audio translation function being triggered. The audio translation function is triggered in at least one of the following ways: waking up the agent through a specified external audio device, the agent's translation function being triggered, or the agent being in a real-time call with the user and in translation mode.

[0088] In some embodiments, the translation module 601 is further configured to: identify the language of the spoken content in the original audio; translate the spoken content in the first language into the second language; and translate the spoken content in the second language into the first language.

[0089] In some embodiments, the interaction module 603 is further configured to: play translated audio in a second language for spoken content in a first language; and play translated audio in a first language for spoken content in a second language.

[0090] In some embodiments, the interaction module 603 is further configured to: receive the user's selection of a target language, the target language including one or more; and play the audio corresponding to the content of the target language in the translation result.

[0091] In some embodiments, the user translation module 601 is further configured to: receive the user's selection of a target language; and translate the non-target language content in the original audio into the target language.

[0092] In some embodiments, the interaction module 603 is further configured to display the translated text of the original audio in response to the original audio originating from outside the user device.

[0093] In some embodiments, the interaction module 603 is further configured to: display the translated text of the original audio in the interaction interface of the agent in response to the agent running in the foreground; and display a translation icon in the interface of the user device in response to the agent running in the background when the original audio comes from outside the user device.

[0094] In some embodiments, the interaction module 603 is further configured to: display an overlay on top of the currently displayed interface when the agent is running in the background; and display the translated text of the latest segment of the original audio in the overlay.

[0095] In some embodiments, the translated text includes multiple paragraphs, each paragraph having a content size less than a threshold. The interaction module 603 is further configured to: in the interaction interface of the agent, in response to the original audio originating from the user device and the user triggering a certain paragraph, play the translated audio corresponding to the paragraph.

[0096] In some embodiments, the interaction module 603 is further configured to display one or more speakers in the original audio, and the translated text of the original audio corresponding to each speaker.

[0097] Figure 7 This diagram illustrates a block diagram of an electronic device according to some embodiments of the present disclosure. Memory 71 is used to store one or more computer-readable instructions. Memory 71 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 71 may, for example, store an operating system, applications, a boot loader, a database, and other programs, as well as various applications and various data.

[0098] The processor 72 is configured to execute computer-readable instructions to implement the method described in any of the foregoing embodiments. Specific implementations of each step of the method can be found in the above embodiments; repeated details will not be elaborated upon here.

[0099] Processor 72 can be configured to execute Figures 1 to 5 The various methods described herein. The processor 72 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be an x87 or ARM architecture, etc.

[0100] The processor 72 and the memory 71 can communicate with each other directly or indirectly. For example, the processor 72 and the memory 71 can communicate via a network. The network can include a wireless network, a wired network, and / or any combination of wireless and wired networks. The processor 72 and the memory 71 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0101] It should be noted that Figure 7 The components of the electronic device 7 shown are merely exemplary and not limiting. The electronic device 7 may have other components depending on the specific application requirements. The processor 72 can control other components in the electronic device 7 to perform desired functions.

[0102] Electronic device 7 can be implemented by software, firmware and / or hardware, and can be integrated into a device with the relevant application installed.

[0103] Figure 8Block diagrams of electronic devices according to other embodiments of the present disclosure are shown. Figure 8 The electronic device 8 shown can be a computer system with a dedicated hardware structure, which can perform corresponding functions when the relevant application is installed.

[0104] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet PCs, portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.

[0105] like Figure 8 As shown, the Central Processing Unit (CPU) 81 performs various processes based on a program stored in the Read-Only Memory (ROM) 82 or a program loaded from the storage section 88 into the Random Access Memory (RAM) 83. The RAM 83 stores data required as needed when the CPU 81 performs various processes, etc. The CPU is merely exemplary; it could also be other types of processors, such as the various processors described above. The ROM 82, RAM 83, and storage section 88 can be various forms of computer-readable storage media. It should be noted that although... Figure 8 The diagram shows ROM 82, RAM 83 and storage section 88, but one or more of them may be combined or located in the same or different memory or storage modules.

[0106] CPU 81, ROM 82 and RAM 83 are interconnected via bus 84. Input / output interface 85 is also connected to bus 84.

[0107] The following components are connected to the input / output interface 85: input section 86, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 87, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 88, including hard disks, magnetic tapes, etc.; and communication section 89, including network interface cards such as LAN cards, modems, etc. The communication section 89 allows communication processing via a network such as the Internet. It is easy to understand that, although... Figure 8 The portion of the electronic device 8 shown communicates via bus 84, but it may also communicate via a network or other means, wherein the network may include a wireless network, a wired network, and / or any combination of wireless and wired networks.

[0108] As needed, drive 810 is also connected to input / output interface 85. Removable media 811, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 810 as needed, so that computer programs read from them can be installed into storage section 88 as needed.

[0109] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 811.

[0110] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to perform the methods described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 89, or installed from storage section 88, or installed from ROM 82. When the computer program is executed by CPU 81, the methods of embodiments of this disclosure are performed.

[0111] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0112] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.

[0113] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Computer instructions are stored on the computer-readable storage medium that, when executed by a processor, implement the methods described in any of the foregoing embodiments.

[0114] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0115] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0116] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the methods described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.

[0117] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0119] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0120] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. An audio translation method, comprising: Translate the raw audio acquired by the user device; Determine whether the original audio source is external or internal to the user device; In response to the original audio originating from outside the user device, the translated audio of the original audio is played by default. Since the original audio originates from within the user device, the translated audio is not played by default; instead, the translated text of the original audio is displayed.

2. The audio translation method according to claim 1 further includes: In response to the audio translation function being triggered, the system listens to external and internal audio of the user device. The audio translation function is triggered in at least one of the following ways: the agent is awakened by a specified external audio device, the agent's translation function is triggered, or the agent is in a real-time call with the user and is in translation mode.

3. The audio translation method according to claim 1, wherein, The translation of the raw audio acquired by the user device includes: Identify the language used in the spoken content of the original audio; Translate spoken content in the first language into the second language; Translate the spoken content in the second language into the first language.

4. The audio translation method according to claim 3, wherein, The translated audio that plays the original audio includes: For the spoken content in the first language, play the translated audio in the second language; For the spoken content in the second language, play the translated audio in the first language.

5. The audio translation method according to claim 3, wherein, The translated audio that plays the original audio includes: Receive the user's selection of a target language, wherein the target language includes one or more; Play the audio corresponding to the content in the target language in the translation results.

6. The audio translation method according to claim 1, wherein, The translation of the raw audio acquired by the user device includes: Receive the user's selection of the target language; Translate the non-target language content in the original audio into the target language.

7. The audio translation method according to claim 1, further comprising: In response to the original audio originating from outside the user device, the translated text of the original audio is displayed.

8. The audio translation method according to claim 1 or 7, wherein, The translated text displayed for the original audio includes: In response to the agent running in the foreground, the translated text of the original audio is displayed in the agent's interactive interface; In response to the agent running in the background, if the original audio originates from outside the user device, a translation icon is displayed on the user device's interface.

9. The audio translation method according to claim 1 or 7, wherein, The translated text displayed for the original audio includes: When the agent is running in the background, a floating layer is displayed on top of the currently displayed interface; The translation of the latest segment of the original audio is displayed in the overlay.

10. The audio translation method according to any one of claims 1 to 7, wherein, The translated text comprises multiple paragraphs, each paragraph having a content size less than a threshold; the translation method further includes: In the interactive interface of the intelligent agent, in response to the original audio originating from the user's device and the user triggering a certain segment, the translated audio corresponding to the segment is played.

11. The audio translation method according to claim 1 or 7, wherein the translated text displayed from the original audio includes: Display one or more speakers from the original audio, and the translated text of the original audio corresponding to each speaker.

12. An electronic device, comprising: Memory; as well as A processor coupled to the memory, the processor being configured to execute the audio translation method as described in any one of claims 1 to 11 based on instructions stored in the memory.

13. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the audio translation method of any one of claims 1 to 11.

14. A computer program product, when run on a computer, causes the computer to implement the audio translation method according to any one of claims 1 to 11.