Method for multimedia data transmission, electronic device, and storage medium
The AI translation continuity function addresses cross-language communication challenges by real-time multimedia data transmission, improving efficiency and continuity in multi-user scenarios.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-08-28
- Publication Date
- 2026-07-23
Smart Images

Figure KR2025013185_23072026_PF_FP_ABST
Abstract
Description
METHOD FOR MULTIMEDIA DATA TRANSMISSION, ELECTRONIC DEVICE, AND STORAGE MEDIUM
[0001] The present disclosure relates to the field of computer technology, and more particularly, to a method and device for multimedia data transmission, an electronic device, a storage medium, and a computer program product.
[0002] With the continuous advancement of globalization, the demand for cross-language communication is growing. In multinational enterprises, there are often situations where colleagues with different linguistic and cultural backgrounds communicate with one another, and at this time, multiple users with different linguistic backgrounds will have demands for content translation.
[0003] For example, multiple users with the same language background may share a document in a language that they all understand via a screen sharing function, and communicate with each other in the language that they all understand. However, when a user with a different language background wants to join the communication, due to the language barrier between the user and the current users, the user cannot understand specific content of the communication in time, which will lead to obstruction of the communication process, thereby seriously affecting the efficiency of communication.
[0004] According to an embodiment of the present disclosure, a method for multimedia data transmission performed by a first device is disclosed. The method may comprise, based on receiving an instruction for enabling multimedia data translation, obtaining target multimedia data by the first device. The method may comprise determining whether the first device is capable of translating the obtained target multimedia data. The method may comprise translating the target multimedia data to obtain translation content when the first device is capable of translating the obtained target multimedia data. The method may comprise transmitting the translation content to at least one other device for displaying the translation content on the at least one other device.
[0005] According to an embodiment of the present disclosure, an electronic device for multimedia data transmission is disclosed. The electronic device may comprise at least one processor and a memory for storing at least one instruction executable by the at least one processor. The at least one processor may be configured to execute the at least one instruction to, based on receiving an instruction for enabling multimedia data translation, obtain target multimedia data. The at least one processor may be configured to execute the at least one instruction to determine whether the electronic device is capable of translating the obtained target multimedia data. The at least one processor may be configured to execute the at least one instruction to translate the target multimedia data to obtain translation content when the electronic device is capable of translating the obtained target multimedia data. The at least one processor may be configured to execute the at least one instruction to transmit the translation content to at least one other device for displaying the translation content on the at least one other device.
[0006] According to an embodiment of the present disclosure, a method for multimedia data transmission applied to a first device performing multimedia data communication with a second device is disclosed. The method may comprise, in response to receiving an instruction for enabling multimedia data translation, acquiring target multimedia data in the first device. The method may comprise, in a case where the first device has a translation capability, translating the target multimedia data to obtain translation content. The method may comprise transmitting the translation content to the second device for display on the second device.
[0007] According to an embodiment of the present disclosure, a method for multimedia data transmission performed by a second device performing multimedia data communication with a first device is disclosed. The method may comprise, in a case where the first device having a translation capability, receiving translation content for target multimedia data obtained by the first device from the first device, wherein the translation content is obtained by translating the target multimedia data by the first terminal in response to receiving an instruction for enabling multimedia data translation. The method may comprise displaying the translation content.
[0008] According to an embodiment of the present disclosure, a device for multimedia data transmission applied to a first device performing multimedia data communication with a second device is disclosed. The device may comprise a multimedia data acquisition module configured to, in response to receiving an instruction for enabling multimedia data translation, acquire target multimedia data in the first device. The device may comprise a translation module configured to, in a case where the first device has a translation capability, translate the target multimedia data to obtain translation content, and transmit the translation content to the second device for display on the second device.
[0009] According to an embodiment of the present disclosure, an electronic device for multimedia data transmission is disclosed. The electronic device may comprise a processor, and a memory for storing instructions executable by the processor. The processor is configured to execute the instructions to implement the method for multimedia data transmission according to the present disclosure.
[0010] According to an embodiment of the present disclosure, there is provided a computer readable storage medium, wherein instructions in the computer readable storage medium, when executed by a processor, enable an electronic device to execute the method for multimedia data transmission according to the present disclosure.
[0011] According to an embodiment of the present disclosure, a device for multimedia data transmission applied to a second device performing multimedia data communication with a first device may comprise a first translation content receiving module configured to, in a case where the first terminal has a translation capability, receive translation content for target multimedia data in the first terminal from the first terminal, wherein the translation content is obtained by translating the target multimedia data by the first terminal in response to receiving an instruction for enabling multimedia data translation. The device may comprise a translation content display module configured to display the translation content.
[0012] According to an embodiment of the present disclosure, there is provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the method for multimedia data transmission according to the present disclosure.
[0013] It should be understood that the above general description and the following detail description are merely exemplary and illustrative without limiting the present disclosure.
[0014] The drawings here are incorporated into the description and constitute a part of the description, illustrate embodiments conforming to the present disclosure, explain a principle of the present disclosure together with the description, and do not constitute inappropriate definition for the present disclosure.
[0015] FIG. 1 is a schematic diagram illustrating a scenario in which multiple users with different linguistic and cultural backgrounds conduct an online conference;
[0016] FIG. 2 is a schematic diagram illustrating a translation continuity architecture according to an embodiment of the present disclosure;
[0017] FIG. 3 is a schematic diagram illustrating a scenario in which a translation continuity function is used in an online conference according to an embodiment of the present disclosure;
[0018] FIG. 4A is a flowchart illustrating a method for multimedia data transmission according to an embodiment of the present disclosure;
[0019] FIG. 4B is a flowchart illustrating a method for multimedia data transmission according to an embodiment of the present disclosure;
[0020] FIG. 5 is a schematic diagram illustrating a continuity floating window popped-up when entering into a particular interface of a third-party App according to an embodiment of the present disclosure;
[0021] FIG. 6 is a schematic diagram illustrating that a plurality of nearby devices searched by a Host device to perform AI translation continuity according to an embodiment of the present disclosure;
[0022] FIG. 7 is a schematic diagram illustrating a flowchart of performing audio translation continuity at a Host terminal according to an embodiment of the present disclosure;
[0023] FIG. 8 is a schematic diagram illustrating a flowchart of performing screen content translation continuity at the Host terminal according to an embodiment of the present disclosure;
[0024] FIG. 9 is a schematic diagram illustrating a flowchart of performing audio translation continuity at a Client terminal according to an embodiment of the present disclosure;
[0025] FIG. 10 is a schematic diagram illustrating a flowchart of performing screen translation continuity at the Client terminal according to an embodiment of the present disclosure;
[0026] FIG. 11 is a schematic diagram illustrating a scenario in which AI translation is continued on a tablet according to an embodiment of the present disclosure;
[0027] FIG. 12 is a schematic diagram illustrating a scenario in which AI translation is continued on a television according to an embodiment of the present disclosure;
[0028] FIG. 13 is a diagram illustrating that only a Host terminal has an AI translation capability according to an embodiment of the present disclosure;
[0029] FIG. 14 is a schematic diagram illustrating an effect of projecting translation content onto a Client device using a projection protocol according to an embodiment of the present disclosure;
[0030] FIG. 15 is a schematic diagram illustrating displaying translated screen content on a second device according to an embodiment of the present disclosure;
[0031] FIG. 16 is a schematic diagram illustrating displaying translated Korean screen content, original Chinese subtitle content and translated Korean subtitle content on the second device according to an embodiment of the present disclosure;
[0032] FIG. 17 is a schematic diagram illustrating displaying original subtitle content and translated subtitle content on the second device according to an embodiment of the present disclosure;
[0033] FIG. 18 is a flowchart illustrating a method for multimedia data transmission according to an embodiment of the present disclosure;
[0034] FIG. 19 is a block diagram illustrating a device for multimedia data transmission according to an embodiment of the present disclosure;
[0035] FIG. 20 is a block diagram illustrating a device for multimedia data transmission according to an embodiment of the present disclosure; and
[0036] FIG. 21 is a block diagram illustrating an electronic device according to an embodiment of the present disclosure.
[0037] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and fully described hereinafter with reference to the drawings.
[0038] The terms used in the present disclosure will be briefly described, and then embodiments of the present disclosure will be described in detail.
[0039] The terms used in the present disclosure are general terms as much as possible and have been widely used in consideration of the functions of the present disclosure. However, they may be changed according to the intent of a person skilled in the art, precedent, the advent of new technologies, or the like. Also, particular cases may include terms arbitrary selected by the applicant, and in such cases, the meaning of the terms will be described in detail in the corresponding description. Therefore, the terms used in the present disclosure should be defined based on the meanings of the terms and the content throughout the present disclosure, rather than simply based on the titles of the terms.
[0040] It may be advantageous to set forth definitions of certain words and phrases used throughout this disclosure. Thus, the terms "transmit," "receive," and "communicate," as well as derivatives thereof, encompass both direct and indirect communication. The terms "include" and "comprise," as well as derivatives thereof, mean inclusion without limitation. The term "or" is inclusive, meaning and / or. The phrase "associated with," as well as derivatives thereof, means to include, be included within, interconnect with, contain, be contained within, connect to or with, couple to or with, be communicable with, cooperate with, interleave, juxtapose, be proximate to, be bound to or with, have, have a property of, have a relationship to or with, or the like.
[0041] Moreover, various functions described below may be implemented or supported by one or more computer programs, each of which is formed from a computer readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, sets of instructions, procedures, functions, objects, classes, instances, related data, or portions thereof adapted for implementation in a suitable computer readable program code. The phrase "computer readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer readable medium" includes any type of medium capable of being accessed by a computer, such as read-only memory (ROM), random access memory (RAM), a hard disk drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. A "non-transitory" computer readable medium excludes wired, wireless, optical, or other communication links that transport transitory electrical or other signals. A non-transitory computer readable medium includes media where data can be permanently stored and media where data can be stored and later overwritten, such as a rewritable optical disc or an erasable memory device.
[0042] As used herein, terms and phrases such as "have," "may have," "include," or "may include" a feature (like a number, function, operation, or component such as a part) indicate the existence of the feature and do not exclude the existence of other features. Also, as used herein, the phrases "A or B," "at least one of A and / or B," or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B," "at least one of A and B," and "at least one of A or B" may indicate any of (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. Further, as used herein, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device (or a first device) and a second user device (or a second device) may indicate different user devices from each other, regardless of the order or importance of the devices. A first component may be referred to as a second component and vice versa without departing from the scope of this disclosure.
[0043] It will be understood that, when an element (such as a first element) is referred to as being (operatively or communicatively) "coupled with / to" or "connected with / to" another element (such as a second element), the element can be coupled or connected with / to the other element directly or via a third element. In contrast, it will be understood that, when an element (such as a first element) is referred to as being "directly coupled with / to" or "directly connected with / to" another element (such as a second element), no other element (such as a third element) intervenes between the element and the other element.
[0044] As used herein, the phrase "configured (or set) to" may be interchangeably used with the phrases "suitable for," "including the capacity to," "designed to," "adapted to," "made to," or "capable of" depending on the circumstances. The phrase "configured (or set) to" does not essentially mean "specifically designed in hardware to." Rather, the phrase "configured to" may indicate that a device can perform an operation together with another device or parts (or component). For example, the phrase "processor configured (or set) to perform A, B, and C" may mean a generic-purpose processor (such as a CPU or application processor) that may perform the operations by executing one or more software programs stored in a memory device or a dedicated processor (such as an embedded processor) for performing the operations.
[0045] The terms and phrases as used herein are provided merely to describe some embodiments of this disclosure but not to limit the scope of other embodiments of this disclosure. It is to be understood that the singular forms "a," "an," and "the" include plural references unless the context clearly dictates otherwise. All terms and phrases, including technical and scientific terms and phrases, used here have the same meanings as commonly understood by one of ordinary skill in the art to which the embodiments of this disclosure belong. It will be further understood that terms and phrases, such as those defined in commonly used dictionaries, should be interpreted as including a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein. In some cases, the terms and phrases defined herein may be interpreted to exclude embodiments of this disclosure.
[0046] Examples of an "electronic device" according to embodiments of this disclosure may include at least one of a smartphone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head-mounted device (HMD), electronic clothes, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch). Other examples of the "electronic device" include a smart home appliance. Examples of the smart home appliance may include at least one of a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave oven, a washer, a dryer, an air cleaner, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC, APPLETV, or GOOGLE TV), a smart speaker or speaker with an integrated digital assistant (such as SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a gaming console (such as an XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camcorder, or an electronic picture frame. Still other examples of the "electronic device" include at least one of various medical devices (such as various portable medical measuring devices (like a blood sugar measuring device, a heartbeat measuring device, or a body temperature measuring device), a magnetic resonance angiography (MRA) device, a magnetic resonance imaging (MRI) device, a computed tomography (CT) device, an imaging device, or an ultrasonic device), a navigation device, a global positioning system (GPS) receiver, an event data recorder (EDR), a flight data recorder (FDR), an automotive infotainment device, a marine electronic device (such as a marine navigation device or a gyro compass), avionics, security devices, vehicular head units, industrial or home robots, automatic teller machines (ATMs), point of sales (POS) devices, or Internet of Things (IoT) devices (such as a bulb, various sensors, electricity or gas meter, sprinkler, fire alarm, thermostat, street light, toaster, fitness equipment, hot water tank, heater, or boiler). Other examples of the "electronic device" include at least one part of a piece of furniture or building / structure, an electronic board, an electronic signature receiving device, a projector, or various measuring devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). Note that, according to various embodiments of this disclosure, the "electronic device" may be one or a combination of the above-listed devices. According to some embodiments of this disclosure, the "electronic device" may be a flexible electronic device. The electronic device disclosed here is not limited to the above-listed devices and may include new electronic devices depending on the development of technology.
[0047] As used herein, the term "user" may refer to a human or an artificial intelligence (AI)-based electronic device. In this disclosure, the expression "and / or" includes a combination of two or more of the described components or any one of the described components.
[0048] Also, the terms "portion," "module," etc. described in this disclosure denote a unit configured to process at least one function or operation, and the "portion," or the "module" may be realized as hardware or software, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC), or a combination of the hardware and the software. The term "portion" used in an embodiment of the present disclosure does not have a meaning limited to software or hardware. A "portion" described in the present disclosure may be configured to be in a storage medium which may be addressed or may be configured to utilize one or more processors. According to an embodiment of the present disclosure, a "portion" may include components, such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, sub-routines, segments of a program code, drivers, firmware, microcode, a circuit, data, a database, data structures, tables, arrays, and variables. Functions provided through a predetermined component or a predetermined "portion" may be combined to reduce the number of functions or may be divided into additional components. Also, according to an embodiment, a "portion" may include one or more processors.
[0049] According to an embodiment of the present disclosure, each of blocks of the flowcharts and combinations of the flowcharts may be performed by computer program instructions. The computer program instructions may be loaded onto a general-purpose computer, a specialized computer, or a processor of another programmable data processing device. The instructions executed by the computer or the processor of the other programmable data processing device may generate a medium for performing the functions described in the flowchart block(s). The computer program instructions may also be stored in a computer-available or computer-readable memory configured for the computer or the other programmable data processing device in order to realize the functions in a predetermined way. The instructions stored in the computer-available or computer-readable memory may also produce manufacturing items embedding an instruction medium for performing the functions described in the flowchart block(s). The computer program instructions may also be loaded onto the computer or the other programmable data processing device.
[0050] Furthermore, each block in the flowcharts may represent a module, segment, or a part of code including one or more executable instructions to perform particular logic function(s). According to an embodiment of the present disclosure, the functions described with respect to the blocks may also be performed not according to an order. For example, two blocks illustrated in succession may be executed substantially concurrently, or the blocks may sometimes be executed in a reverse order, depending on the functions involved therein.
[0051] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the accompanying drawings, so that the embodiments of the present disclosure may be easily implemented by one of ordinary skill in the art. However, the embodiments of the present disclosure may have different forms and should not be construed as being limited to the embodiments of the present disclosure described herein. Also, in the drawings, parts not related to descriptions are omitted for clarity of the description of the embodiments of the present disclosure, and throughout the specification, like reference numerals are used for like elements. With the continuous advancement of globalization, the demand for cross-language communication is growing. The level of development of translation technology, as a vital tool for bridging cultures and fostering international collaboration, directly impacts efficiency and quality of global information flow. Although current translation tools can provide basic translation services, in a multi-person conference scenario, when there are people with different language and cultural backgrounds attending the conference, a separate translation tool may not satisfy the needs of synchronous communication between multiple users with different language and cultural backgrounds attending the conference.
[0052] The present disclosure provides a method and device for multimedia data transmission, an electronic device, a storage medium and a computer program product, so as to at least address the problem of obstruction of the communication process and serious degradation of communication efficiency when multiple users do not speak the same language in the related technologies.
[0053] FIG. 1 is a schematic diagram illustrating a scenario in which multiple users with different linguistic and cultural backgrounds conduct an online conference. Referring to FIG. 1, in an online video conference, a Chinese colleague A 10 and a remote Chinese client C 20 share a Chinese document through a conference screen-sharing function and communicate about the document in Chinese. During the video conference, they encounter a problem that needs to be confirmed together with a Korean colleague B 30. However, Chinese of Korean colleague B 30 is not good enough to understand the content of the shared Chinese document as well as the content of the communication between the Chinese colleague A 10 and the Chinese client C 20. This will lead to obstruction of the communication process, and seriously affect the efficiency of the communication.
[0054] A method and device for multimedia data transmission, an electronic device, a storage medium, and a computer program product provided by the present disclosure may, in a multi-user communication scenario, enable an AI translation continuity function if multiple users do not speak the same language as each other, so that the corresponding translation content can be displayed on a second device 200 (e.g., a client device) that requires translation while communicating through a first device 100 (e.g., a host device). In this way, a user that requires translation may understand the specific content of the communication in time to further ensure smooth progress of the communication process, thereby improving the efficiency of the communication. For example, the first device 100 may be the device of the Chinese colleague A 10 shown in FIG. 1, and the second device 200 may be the device of the Korean colleague B 30 shown in FIG. 1. The first device 100 may be referred to as a host device, a host terminal, an electronic device, or a first terminal. The second device 200 may be referred to as a client device, a client terminal, an electronic device, or a second terminal.
[0055] FIG. 2 is a schematic diagram illustrating a translation continuity architecture according to an embodiment of the present disclosure.
[0056] Referring to FIG. 2, the first device 100 and the second device 200 are illustrated. The first device 100 may be, but is not limited to, a smart phone, a tablet, etc.; and the second device 200 may be, but is not limited to, a smart phone, a tablet, a Personal Computer (PC), a Television (TV), etc.
[0057] The first device 100 may include a first VCC (Video Call Continuity) app 110, a first audio HAL (Hardware Abstraction Layer) 120, and a third-party app 130 which may be a video call type application. A VCC app is an app for transferring a video call from a device (e.g., a phone) to another device (e.g., a Tablet or TV). Accordingly, the first VCC app 110 may transfer a video call from the first device 100 to the second device 200. The first VCC app 110 may include a first screen content translation module 111(screen translation / Optical Character Recognition (OCR)), and a first audio transfer module 112. The first audio HAL may include a first voice recognition module (e.g., Voice Over IP Automatic Speech Recognition (VOIP ASR)). The configuration of the first device 100 is not limited to that illustrated in FIG. 2. For example, the first device 100 may be configured as the electronic device 2100 shown in FIG. 21. When the first device 100 is configured as the electronic device 2100, operations performed based on the components (e.g., the first VCC (Video Call Continuity) app 110, the first audio HAL (Hardware Abstraction Layer) 120, and the third-party app 130) included in the first device 100 may be executed based on the memory 2101 and the processor 2102.
[0058] The second device 200 may include a second VCC app 210 and a second audio HAL 220. The second VCC app 210 may receive a video call transferred from the first device 100 and perform video call continuity on the second device 200. The second VCC app 210 may include a second screen content translation module 211, and a second audio transfer module 212. The second Audio HAL 221 may include a second voice recognition module 221. The configuration of the second device 200 is not limited to that illustrated in FIG. 2. For example, the second device 200 may be configured as the electronic device 2100 shown in FIG. 21. When the second device 200 is configured as the electronic device 2100, operations performed based on the components (e.g., the second VCC (Video Call Continuity) app 210, and the second audio HAL (Hardware Abstraction Layer) 220) included in the second device 200 may be executed based on the memory 2101 and the processor 2102. The second device 200 may configured as shown in FIG. 13. The first VCC app 110 may be used for providing a continuity function to the third-party app 130 (or a third-party video call apps).
[0059] The first screen content translation module 111 may be used for controlling OCR recognition and controlling the translation logic. That is, the first screen content translation module 111 may utilize an OCR technique to perform character recognition on screen content that is displayed on the first device 100. Then, the first screen content translation module 111 may translate the recognized character content by utilizing OCR. Then, the first screen content translation module 111 may erase the original screen content by using OpenCV (Open Source Computer Vision Library), and may fill the erased positions with the translated content. The first screen content translation module 111 may transfer the translated content to the second device 200 via a data bus 150.
[0060] The first audio transfer module 112 may be used for implementing an AI audio translation control logic, and may specifically include the first device 100 requesting to query an AI audio translation capability of the second device 200, the first device 100 transferring a result of the AI audio translation to the second device 200, and the first device 100 controlling a language switching capability of the AI audio translation of the second device 200, etc.
[0061] The first audio HAL 120 may be used to implement video ASR and a translation function, and may provide a related interface that can be called by the first VCC app 110. The term "video ASR" may refer to a technology that automatically recognizes and converts spoken language into text during real-time video calls or online conferences. This technology is typically used to support multilingual communication, enable live transcription, or assist translation systems during video-based interactions (e.g., a video call or an online conference).
[0062] The voice recognition module 121 may be included in the audio HAL 120 and may be configured to perform automatic speech recognition (ASR) based on input voice data received during real-time video calls or online conferences. The voice recognition module 121 may recognize spoken language from the input audio signals and convert it into text data. The recognized text may then be used by the first VCC app 110 for translation, live captioning, or other multimedia processing tasks. The voice recognition module 121 may operate under the control of the first VCC app 110 or the audio transfer module 112 via a corresponding API exposed by the audio HAL 120.
[0063] The third-party app 130 may provide video call functions and may perform the continuity function by using the first VCC app 110. That is, when the third-party app 130 enters a video call interface or an online conference interface, the third-party app 130 may display an option of a video continuity floating window on the first device 100. When a user may click the option, the third-party app 130 may perform the continuity function by invoking the first VCC app 110. The continuity function may include a function that allows a user to transfer the ongoing video call to another device (e.g., the second device 200) and continue the session seamlessly by invoking the first VCC app 110. The data bus 150 may be used for data transfer between the first device 100 and the second device 200. For example, an AI translation result of the first device 100 may be transmitted to the second device 200 through the data bus 150.
[0064] A control bus 151 may be used for exchanging basic information and transmitting control commands. For example, the first device 100 may transmit, via the control bus 151, a request to the second device 200 to query whether the second device 200 has an AI audio translation capability and / or AI screen content translation capability, and the second device 200 may return information related to the AI audio translation capability and / or AI screen content translation capability to the first device 100 via the control bus 151. The first device 100 may transmit a language setting command for the AI audio translation to the second device 200 via the control bus 151c.
[0065] The second device 200 may include a second VCC app 210, which may include components such as a second screen content translation module 211 and a second audio transfer module 212. The second device 200 may also include a second audio HAL 220, which may include components such as a second voice recognition module 221. These components included in the second device 200 may operate based on data and control information received from the first device 100 via the data bus 150 and the control bus 151.
[0066] For example, when screen content translated into a target language is received from the first device 100, the second screen content translation module 211 may display the translated content on the screen of the second device 200 based on the received data. When audio data translated into the target language is received from the first device 100, the second audio transfer module 212 may reproduce and output the translated audio to the user of the second device 200. The target language may refer to a language into which source content, such as screen content or audio data, is translated by a translation module (e.g., the first screen content translation module 111, the first audio transfer module 112 including audio translation control logic or performing an audio translation function, the second screen content translation module 211, the second audio transfer module 212 including audio translation control logic or performing an audio translation function).
[0067] In addition, when untranslated original content is received from the first device 100, the second screen content translation module 211 may perform translation into the target language within the second device 200 and display the translated content on the second device 200. The second audio transfer module 212 may translate the received original audio data into the target language and output the translated audio via the second device 200.
[0068] The second voice recognition module 221 may work in conjunction with the second audio transfer module 212 to recognize voice signals received from the first device 100, or recognize voice input from the user of the second device 200.
[0069] In addition, when a query is received from the first device 100 via the control bus 151 as to whether the second device 200 has an AI audio translation capability or a screen content translation capability, the second VCC app 210 may check the status of the relevant modules included in the second device 200, construct information indicating the availability of such capabilities, and transmit a response including information related to the AI audio translation capability and / or AI screen content translation capability to the first device 100 via the control bus 151.
[0070] In the present disclosure, only the first device 100 may include the screen translation / OCR module and the VOIP ASR module. Alternatively, only the second device 200 may include the screen translation / OCR module and the VOIP ASR module. Alternatively, the first device 100 and the second device 200 each may include the screen translation / OCR module and the VOIP ASR module. The present disclosure does not make a specific definition thereof, and the previous embodiments are only for exemplary illustration.
[0071] FIG. 3 is a schematic diagram illustrating a scenario in which a translation continuity function is used in an online conference according to an embodiment of the present disclosure.
[0072] Referring to FIG. 3, Chinese colleague A 10 may synchronize the content of the document on his / her device (the first device 100) to the device (the second device 200) of Korean colleague B 20. At this time, the first device 100 of Chinese colleague A 10 may still display the Chinese document, and Chinese colleague A 10 may still use Chinese to communicate with Chinese client C 30, the second device 200 of Korean colleague B 20 may display the content of the Korean document obtained by translating the Chinese document, and the second device 200 of Korean colleague B 20 may further display original dialogue subtitle between Chinese colleague A 10 and Chinese client C 30, as well as the translation subtitle. In this way, even if Korean colleague B 20 does not understand Chinese, he / she may also learn the specific content of the communication between the parties participating in the conference in time, which may ensure the smooth progress of the communication process and improve the efficiency of communication.
[0073] In the present disclosure, smooth communication in a multi-person conference with participants having different language backgrounds may be achieved by combining a video conference application continuity function between multiple devices with the AI translation function.
[0074] FIG. 4A is a flowchart illustrating a method for multimedia data transmission according to an embodiment of the present disclosure. The method for multimedia data transmission may be applied to, or performed by, the first device 100. The first device 100 may refer to a host device, a first terminal, an electronic device, or a first electronic device. The host device 100 and the second device 200 may perform multimedia data communication. The second device 200 may refer to a client device, a second terminal, an electronic device, or a second electronic device.
[0075] Referring to FIG. 4A, in operation 401, based on receiving an instruction for enabling multimedia data translation, the first device 100 may obtain or acquire target multimedia data at the first device 100. In operation 401, the first device 100 may obtain or acquire target multimedia data at the first device 100, in response to receiving the instruction for enabling multimedia data translation.
[0076] For example, it may be pre-configured that when a user enters an interface of third-party app 130, the first device 100 may be controlled to pop up a continuity floating window on the interface by using the first VCC app 110. The interface may include, but is not limited to, a video call interface, an online conference interface, or an audio conference interface, etc.
[0077] FIG. 5 is a schematic diagram illustrating a continuity floating window that is popped up when entering an interface of the third-party app 130 according to an embodiment of the present disclosure.
[0078] For example, the user (e.g., the Chinese colleague A 10) may enable the AI translation continuity function by clicking the popped-up continuity floating window 501 in FIG. 5. For example, the first device 100 may search for surrounding devices through BLE or Wi-Fi P2P when the user clicks the popped-up continuity floating window 501 in FIG. 5. Then, the user may select a device from a plurality of searched devices, and further enable the AI translation continuity function. The user may change the target language of translation through setting.
[0079] FIG. 6 is a schematic diagram illustrating a plurality of nearby devices searched by the first device 100 to perform AI translation continuity according to an embodiment of the present disclosure. The first device 100 may display information (e.g., YY TV 610, YY's Tablet 620) related to AI translation continuity target device selection shown in FIG. 6 through a separate window 600, but is not limited thereto.
[0080] According to an embodiment of the present disclosure, the target multimedia data may include screen content of the first device 100 and / or subtitle content obtained by performing speech recognition on audio content of the first device 100. The audio content may be participants' speech content during communication, or may also be audio files referred to in shared video materials; the screen content may include shared document content, video files, PPT content, and the like.
[0081] In operation 402, in a case where the first device 100 has a translation capability (or translation capabilities), the first device 100 may translate the target multimedia data to obtain translation content, and may transmit the translation content to the second device 200 (or at least one other device) for display on the second device 200 (or on the at least one other device). The at least one other device may include at least one device selected based on the information related to the AI translation continuity target device selection shown in FIG. 6. For example, the first VCC app 110 of the first device 100 may perform translation continuity according to the AI translation capabilities of the first device 100 and the second device 200, including translation continuity of the audio content and / or translation continuity of the screen content.
[0082] For example, when a user is in an online conference (e.g., a video conference or a voice conference), the first device 100 may be a participant in the conference, while the second device 200 may not directly participate in the conference. After the AI translation continuity function is enabled, screen content shared on the first device 100 and the participants' speech content may be translated into a target (or specified) language and displayed on the second device 200. In addition, the translation operation may be performed on either the first device 100 or on the second device 200, depending on whether the first device 100 and the second device 200 support ASR and OCR capabilities.
[0083] According to an embodiment of the present disclosure, the target multimedia data may include audio content. Speech recognition may be performed on the audio content, and for example, ASR technology may be used to perform speech recognition on the audio content to obtain subtitle content, i.e., text content, corresponding to the audio content. Then, the text content may be translated into text content in the target language by the first device 100, the second device 200, or both.
[0084] In this way, by translating the audio content generated during the conference into text content in a target (or specified) language, it becomes convenient for a user of the second device 200 who does not speak the same language to understand in real time the specific content of the communication between the users participating in the conference, thereby further ensuring the smooth progress of the conference process and improving the efficiency of the communication.
[0085] According to an embodiment of the present disclosure, it may further be possible to query whether the first terminal and / or the second terminal has the translation capability. Then, a main device for translating the target multimedia data may be determined according to a query result.
[0086] According to an embodiment of the present disclosure, the first device 100 may transmit a query request to the second device 200 to determine whether the second device 200 has a translation capability. Then, the first device 100 may receive a response including information related to the translation capability of the second device 200 from the second device 200.
[0087] According to an embodiment of the present disclosure, in a case where the first device 100 has the translation capability, the first device 100 may translate the target multimedia data to obtain translation content.
[0088] Regarding an audio content part, whether the first device 100 or the second device 200 may be used to perform translation continuity of the audio content may be determined based on whether the first device 100 or the second device 200 supports ASR capability. For example, if both the first device 100 and the second device 200 have ASR capability, the first device 100 may preferably perform ASR and translation; and if only the first device 100 has ASR capability and the second device 200 does not have ASR capability, the first device 100 may be used to perform ASR and translation.
[0089] FIG. 4B is a flowchart illustrating a method for multimedia data translation according to an embodiment of the present disclosure.
[0090] In operation 403, the first device 100 may obtain or acquire target multimedia data at the first device 100 based on, or in response to, receiving an instruction for enabling multimedia data translation. For example, the first device 100 may receive the instruction to initiate or activate multimedia translation (e.g., from a user input, application event, or a system trigger). Upon receiving the instruction, the first device 100 may obtain, acquire, or capture the relevant target multimedia data, such as screen content or audio content, which may be shared, transmitted, or displayed during a communication session, such as a video conference or online meeting.
[0091] In operation 404, the first device 100 may determine whether it is capable of translating the obtained target multimedia data. For example, the first device 100 may evaluate or determine whether it has internal translation capabilities, such as OCR (Optical Character Recognition) for text on the screen or ASR (Automatic Speech Recognition) for audio speech. Based on this evaluation or determination, the first device 100 may decide whether the translation will be performed by itself or by another device (e.g., the second device 200).
[0092] In operation 405, the first device 100 may translate the target multimedia data to obtain translation content when the first device 100 is capable of translating the obtained target multimedia data. For example, when the first device 100 has translation capability and the target multimedia data includes audio content, the first device 100 may use Automatic Speech Recognition (ASR) to extract subtitle (text) content corresponding to the audio content, and may translate the subtitle content into a target language. For example, when the first device 100 has translation capability and the target multimedia data includes screen content, the first device 100 may use Optical Character Recognition (OCR) to extract text from the screen content, and may translate the extracted text into a target language.
[0093] In operation 406, the first device 100 may transmit the translation content to at least one other device (e.g., the second device 200 such ad a client device) for displaying the translation content on the screen of the at least one other device. The first device 100 may display the translation content on the target device's screen for the user. The at least one other device may correspond to at least one device selected based on AI translation continuity settings, such as those shown in FIG. 6.
[0094] FIG. 7 is a schematic diagram illustrating a flowchart of performing audio translation continuity at a first device according to an embodiment of the present disclosure.
[0095] Referring to FIG. 7, in operation 701, the first device 100 may obtain audio content output by a speaker and audio content collected by a microphone, respectively, after the first device 100 enables the audio translation continuity function. For example, the audio content output by the speaker may include speech content of another user participating in the conference and audio file content referenced in shared video materials. The audio content collected by the microphone may include speech content of the user of the first device 100 and / or speech content of a user of the second device 200.
[0096] In operation 702, the first device 100 may perform speech recognition ASR on the audio content output by the speaker and the audio content collected by the microphone, respectively, to obtain text content. The audio content collected by a microphone of the first device 100 may include the speech of the user at the first device 100 and / or the speech of the user at the second device 200, and the users may speak different languages. In addition, the speech recognition ASR performed by the first device 100 may have multi-language recognition capability and may automatically recognize multi-language audio to further obtain corresponding text content.
[0097] In operation 703, the first device 100 may translate the text content obtained by performing the speech recognition into text content in a target (or specified) language. For example, when the speech content of the user at the first device 100 is collected by the microphone, the speech content of the user at the first device 100 may be translated into text content in a language understandable to the user at the second device 200. For example, when the speech content of the user at the second device 200 is collected by the microphone, the speech content of the user at the second device 200 may be translated into text content in a language understandable to the user at the first device 100. In this way, all participants can understand the speech content of other parties, thereby ensuring the smooth progress of the conference.
[0098] In operation 704, the first device 100 may transmit both the original text content obtained by performing the speech recognition and the text content in the target language obtained by translating the original text content together to the second device 200.
[0099] In operation 705, the second device 200 may display both the original text content obtained by performing the speech recognition and the text content in the target language obtained by translating the original text content. For example, the original speech content of the user of the first device 100 and corresponding translation content, the original speech content of the user of the second device 200 and the corresponding translation content, the original speech content of a counterpart user participating in the conference and corresponding translation content, and the original text content corresponding to audio files referred to by the shared video materials and the corresponding translation content may be displayed in different areas on a screen of the second device 200, respectively, to avoid interference among the respective speech recognition contents from different audio sources.
[0100] For the screen content part, before performing the translation continuity on the screen content, the first device 100 may transmit an instruction to the second device 200 to query whether the second device 200 has OCR and translation capabilities. Then, whether the first device 100 or the second device 200 is used to perform screen translation continuity may be determined depending on whether the first device 100 and the second device 200 support OCR and translation capabilities. For example, when the first device 100 and the second device 200 both support OCR and translation capabilities, or when only the first device 100 has OCR and translation capabilities while the second device 200 does not have OCR and translation capabilities, the OCR and translation capabilities of the first device 100 may be invoked to perform character recognition and translation on the screen content currently displayed on the first device 100, and the translated content may be transmitted to the second device 200 for display through the first VCC app 110 and / or the second VCC app 210.
[0101] FIG. 8 is a schematic diagram illustrating a flowchart of performing screen content translation continuity at the first device according to an embodiment of the present disclosure.
[0102] Referring to FIG. 8, in operation 801, the first device 100 may obtain or acquire a video stream of a screen of the first device 100 through a virtual screen, and may capture a bitmap of a current frame from the video stream at 2-second intervals. By using the virtual screen, the first device 100 may determine whether a change in the bitmap occurs without affecting the user's viewing experience of the actual screen. Then, the first device 100 may compare two consecutively captured bitmaps to determine whether there is a change. When no change in the bitmaps is detected, the first device 100 may return to operation 801. When a change in the bitmaps is detected, the first device 100 may perform operation 802.
[0103] In operation 802, when a change in the bitmaps is detected, the change indicates that the currently displayed video frame-rendered from the video stream-has been updated, i.e., the screen of the first device 100 is transitioning to display new content (i.e., a new video frame). At this time, the first device 100 may trigger an OCR instruction to extract text content requiring translation from the updated video frame.
[0104] In operation 803, the first device 100 may translate the character content extracted from the video frame into character content in a target language.
[0105] In operation 804, the first device 100 may use the OpenCV to erase or remove original character content from the video frame and may fill the corresponding erased or removed positions (or regions) of the video frame with the translated character content in the target language. The first device 100 may save the video frame in which the character content has been replaced.
[0106] In operation 805, the first device 100 may transmit the translated video frame to the second device 200. In the present disclosure, the first device 100 may transmit only the translated screen content image to the second device 200, without transmitting the original video stream, but is not limited thereto. In operation 805, the first device 100 may transmit both the translated screen content image and the original video stream (or the original screen content image) to the second device 200.
[0107] In operation 806, the second device 200 may display the translated video frame. If the second device 200 receive both the translated screen content image and the original video stream, the second device 200 may display both the translated screen content image and the original video stream in a manner that allows the user of the second device 200 to distinguish between the translated screen content image and the original video stream.
[0108] According to an embodiment of the present disclosure, in a case where the first device 100 does not have the translation capability and the second device 200 has the translation capability, the first device 100 may transmit the target multimedia data and a translation instruction to the second device 200, and the second device 200 may translate the target multimedia data based on the received the translation instruction.
[0109] For the audio content part, if the first device 100 does not have the ASR capability, and only the second device 200 has the ASR capability, the first device 100 may transmit the audio content of the first device 100 to the second device 200 through the first VCC app 110, so that the second device 200 may perform ASR and translation using the second VCC app 210 to obtain or acquire original text content corresponding to the audio content and text content in a target (or specified) language obtained by translating the original text content.
[0110] FIG. 9 is a schematic diagram illustrating a flowchart for performing audio translation continuity at the second device 200 according to an embodiment of the present disclosure.
[0111] Referring to FIG. 9, in operation 901, the first device 100 may transmit an audio stream of the first device 100 to the second device 200 through the first VCC app 110.
[0112] In operation 902, in a case where the second device 200 enables a "Speech recognition and translation" function, the second device 200 may respectively extract an audio stream output by a speaker of the first device 100 and an audio stream collected by a microphone of the first device.
[0113] In operation 903, the first device 100 may perform ASR on the audio stream output by the speaker and the audio stream collected by the microphone, respectively, to obtain original text content.
[0114] In operation 904, the first device 100 may translate the original text content recognized through the speech recognition to obtain text content in a target or specified language.
[0115] In operation 905, the second device 200 may display both the original text content corresponding to the audio stream output by the speaker and its translated content, and the original text content corresponding to the audio stream collected by the microphone and its translated content on a screen of the second device 200.
[0116] For the screen content part, if the first device 100 does not support the OCR and translation capability, and only the second device 200 supports the OCR and translation capability, the second device 200 may invoke the OCR and translation capability of the second device 200 and perform character recognition and translation on a video frame transmitted by the first device 100.
[0117] FIG. 10 is a schematic diagram illustrating a flowchart of performing screen translation continuity at the second device 200 according to an embodiment of the present disclosure.
[0118] Referring to FIG. 10, in operation 1001, the first device 100 may transmit the video stream to the second device 200 according to a user input or a user command. The user input or the user command may include a request for screen translation.
[0119] In operation 1002, the second device 200 may capture a bitmap of a current frame of the video stream received from the first device 100 at 2-second intervals.
[0120] In operation 1003, the second device 200 may compare two consecutively captured bitmaps to determine whether a change has occurred between the two consecutively captured bitmaps, in order to determine whether the screen content displayed by the first device 100 has been updated. In a case where the bitmaps do not change, the process may return to operation 1003; and in a case where the bitmaps change, operation 1004 may be performed. The two consecutively captured bitmaps may correspond to two adjacent captured bitmaps or two adjacent intercepted bitmaps.
[0121] In operation 1004, if the bitmaps change, this may indicates that a video frame switch has occurred, i.e., the screen of the first device 100 is currently switching to display a new video frame. At this time, the second device 200 may trigger an OCR instruction to extract character content requiring translation from the current video frame.
[0122] In operation 1005, the second device 200 may translate the character content extracted from the video frame into character content in a target language.
[0123] In operation 1006, the second device 200 may use the OpenCV to erase or remove original character content from the video frame, fill the corresponding erased or removed positions of the video frame with character content in the target language obtained by translating the original character content, and obtain a translated video frame.
[0124] In operation 1007, the second device 200 may display the translated video frame on a screen (or a display) of the second device 200.
[0125] According to an embodiment of the present disclosure, in a case where the first device 100 and the second device 200 do not have the translation capability, the target multimedia data and a translation instruction may be transmitted to a third device 300 (e.g., a cloud server) having a translation capability, wherein the translation instruction is used to instruct the third device 300 to translate the target multimedia data, and the obtained translation content may be transmitted to the second device 200. The third device 300 may be, but is not limited to, a server. For example, the translated content may be transmitted directly from the third device 300 to the second device 200, or the translated content may first be transmitted from the third device 300 to the first device 100, and then from the first device 100 to the second device 200. The present disclosure does not impose any specific limitation in this regard.
[0126] In this way, in a case where the first device 100 and the second device 200 do not have the translation capability, the third device 300 having the translation capability may be enabled to perform translation continuity, thereby avoiding a situation where the communication is forced to be interrupted since the first device 100 and the second device 200 do not have the translation capability, that is, ensuring the smooth communication progress and improving communication efficiency.
[0127] FIG. 11 is a schematic diagram illustrating an effect of continuing AI translation to a tablet according to an embodiment of the present disclosure.
[0128] Referring to FIG. 11, if the second device 200 is a mobile electronic device such as a tablet or the like, since a screen of such a mobile electronic device is relatively smaller than a screen of the first device 100, an entire screen of the second device 200 may be set to display the same proportion of screen content as the first device 100, and the translated screen content may be displayed at a corresponding position of the video frame by replacing the original text content; and results of speech recognition and translation may be displayed in the form of a floating window 1100. In addition, the second device 200 may display an originally used language and a target language of translation. For example, the originally used language may be Chinese, the target language of translation may be Korean, etc.
[0129] Thus, when the screen of the second device 200 is relatively smaller than the screen of the first device 100, the second device 200 may control an audio translation result to be displayed in the form of floating window 1100 that is floatingly displayed above the translated screen content, so that the screen space of the second device 200 can be efficiently utilized and the comprehensiveness and clarity of information display can be ensured.
[0130] FIG. 12 is a schematic diagram illustrating an effect of continuing AI translation to a television according to an embodiment of the present disclosure.
[0131] Referring to FIG. 12, if the second device 200 is an electronic device such as a TV or the like, since a screen of such an electronic device is relatively larger than the screen of the first device 100, an entire screen of the second device 200 may be divided into two areas, one of the areas may be utilized to display the translated screen content in the same screen proportion as the first device 100, and the translated screen content may replace the original text content to be displayed at a corresponding position in the video frame; and the other area may be used to display the results of speech recognition and translation.
[0132] Thus, when the screen of the second device 200 is relatively larger than the screen of the first device 100, the translation continuity result of the audio and the translation continuity result of the screen content may be displayed separately by means of screen division, which may enable effective utilization of the screen space of the second device 200 and ensure the comprehensiveness and clarity of information display.
[0133] According to an embodiment of the present disclosure, the translation content may be projected to the second device 200. For example, the second device 200 may not have the AI translation capability or may not have the second VCC app 210 installed, and at this time, the target multimedia data needs to be translated by using the AI translation capability of the first device 100, and the first device 100 may transfer the translated content to the second device 200 by using a projection protocol. For example, the projection protocol may be but is not limited to, a Miracast protocol.
[0134] FIG. 13 is a diagram illustrating that only a first device 100 has the AI translation capability according to an embodiment of the present disclosure. Referring to FIG. 13, the audio content and the screen content are all translated at the first device 100, and the first device 100 may project the translated content to the second device 200 for display through the Miracast protocol.
[0135] FIG. 14 is a schematic diagram illustrating an effect of projecting translated screen content onto the second device 200 using a projection protocol according to an embodiment of the present disclosure.
[0136] Referring to FIG. 14, a Smart View function may support the Miracast protocol, thus, the translated screen content of the first device 100 may be projected to the second device 200 for display via a Smart View App. At this time, the second device 200 may simultaneously display the original screen content, the translated screen content obtained by translating the original screen content, original subtitles obtained by performing speech recognition on the audio content, and translated subtitles of the original subtitle. In addition, the user may also modify the layout of the translated audio content and the translated video frame content on the screen of the second device 200 according to the actual needs.
[0137] According to an embodiment of the present disclosure, original content may be displayed on the first device 100, wherein the original content may include screen content and / or subtitle content, and translation content corresponding to the original content may be display on the second device 100. FIG. 15 is a schematic diagram illustrating displaying translated screen content on the second device 200 according to an embodiment of the present disclosure. Referring to FIG. 15, the second device 200 may display Korean screen content obtained by translating original Chinese screen content.
[0138] According to an embodiment of the present disclosure, original content may be displayed on the first device 100, wherein the original content may include screen content and / or subtitle content, and the original content may be transmitted to the second device 200, wherein the original content and translation content corresponding to the original content may be display together on the second device 200. FIG. 16 is a schematic diagram illustrating displaying translated Korean screen content, original Chinese subtitle content and translated Korean subtitle content together on the second device 200 according to an embodiment of the present disclosure.
[0139] According to an embodiment of the present disclosure, in a case where the original content includes the subtitle content, the subtitle content may not be displayed on the first device 100, and the subtitle content and the translation content corresponding to the subtitle content may be displayed on the second device 200. FIG. 17 is a schematic diagram illustrating displaying original subtitle content and translated subtitle content on the second device 200 according to an embodiment of the present disclosure. Referring to FIG. 17, the second device 200 may display original Chinese subtitle content and Korean subtitle content obtained by translating the original Chinese subtitle content.
[0140] According to an embodiment of the present disclosure, original content and translation content corresponding to the original content may be displayed on the first device 100, wherein the original content may include screen content and / or subtitle content, and the original content may be transmitted to the second device 200, and the second device 200 may display the original content and the translation content corresponding to the original content.
[0141] According to an embodiment of the present disclosure, when a participant wants to end the translation continuity during the online conference, he or she may click an End button of the translation continuity, and then the first device 100 and the second device 200 may be disconnected and thus the translation continuity may be stopped.
[0142] FIG. 18 is a flowchart illustrating a method for multimedia data transmission according to an embodiment of the present disclosure. The method for multimedia data transmission may be applied to, or performed by, the second device 200, and the second device 200 may perform multimedia data communication with the first device 100.
[0143] Referring to FIG. 18, in operation 1801, when the first device 100 has a translation capability, the second device 200 may receive translation content for target multimedia data from the first device 100. The translation content may be obtained by the first device 100 by translating the target multimedia data in response to, or based on, an instruction to enable multimedia data translation.
[0144] According to an embodiment of the present disclosure, the target multimedia data may include screen content displayed on the first device 100 and / or subtitle content obtained by performing speech recognition on the audio content of the first device 100. The audio content may be participants' speech content during a communication session, or may also be audio files included in shared video materials; the screen content may be shared document content, video files, Power Point Presentations (e.g., PPT content), or other visual materials.
[0145] According to an embodiment of the present disclosure, the second device 200 may receive a translation capability query request from the first device 100. In response to, or based on, the translation capability query request, the second device 200 may transmit information related to the translation capability of the second device 200 to the first device 100.
[0146] According to an embodiment of the present disclosure, when the first device 100 does not have the translation capability and the second device 200 has the translation capability, the second device 200 may receive the target multimedia data and a translation instruction from the first device 100. In response to the received translation instruction or based on the received translation instruction, the second device 200 may translate the target multimedia data to obtain the translation content.
[0147] According to an embodiment of the present disclosure, when the first device 100 and the second device 200 do not have the translation capability, the second device 200 may receive the translation content for the target multimedia data from the third device 300 (e.g., a server). The translation content may be obtained by translating the target multimedia data by the third device 300 having the translation capability. The third device 300 may be, but is not limited to, the server. For example, the second device 200 may receive the translated content transmitted by the third device 300. For example, the first device 100 may receive the translated content transmitted by the third device 300, and then the second device 200 may receive the translated content transmitted by the first device 100. The present disclosure does not make any specific definition thereon.
[0148] In this way, when the first device 100 and the second device 200 do not have the translation capability, the third device 300 having the translation capability may perform translation continuity, thereby avoiding a phenomenon where the communication is forced to be interrupted because the first device 100 and the second device 200 both do not have the translation capability, and thus ensuring the smooth progress of the communication and improving the efficiency of the communication.
[0149] According to an embodiment of the present disclosure, the second device 200 may receive translation content projected from the first device 100. For example, the second device 200 may not have the AI translation capability or may not have the second VCC app210 installed, and at this time, the target multimedia data needs to be translated using the AI translation capability of the first device 100. The first device 100 may transfer the translated content to the second device 200 by using a projection protocol. For example, the projection protocol may be, but is not limited to, a Miracast protocol.
[0150] In step 1802, the second device 200 may display the translation content.
[0151] According to an embodiment of the present disclosure, the second device 200 may display original content and translation content corresponding to the original content on the screen (or display) of the second device 200, wherein the original content may include screen content and / or subtitle content. The original content may further be displayed on the screen (or display) of the first device 100.
[0152] According to an embodiment of the present disclosure, the second device 200 may receive original content from the first device 100, wherein the original content may include screen content and / or subtitle content, and the original content may further be displayed on the first device 100. The second device 200 may display the original content and the translation content corresponding to the original content on the screen (or display) of the second device 200.
[0153] According to an embodiment of the present disclosure, when the original content includes the subtitle content, the second device 200 may display the subtitle content and the translation content corresponding to the subtitle content on the screen (or display) of second device 200, wherein the subtitle content may not be displayed on the first device 100.
[0154] According to an embodiment of the present disclosure, original content and translation content corresponding to the original content may be displayed on the first device 100, wherein the original content may include screen content and / or subtitle content. The second device 200 may receive from the first device 100, and then the second device 200 may display the original content and the translation content corresponding to the original content on the screen (or display) the second device 200.
[0155] FIG. 19 is a block diagram illustrating a device for multimedia data transmission according to an embodiment of the present disclosure.
[0156] Referring to FIG. 19, the device 1900 for multimedia data transmission may include a multimedia data obtaining module 1901 and a translation module 1902. The multimedia data obtaining module 1901 may be referred to as a multimedia data acquisition module.
[0157] The multimedia data obtaining module 1901 may, in response to receiving an instruction for enabling multimedia data translation, obtain or acquire target multimedia data at the first device 100.
[0158] According to an embodiment of the present disclosure, the multimedia data obtaining module 1901 may obtain or acquire the target multimedia data, which may include screen content of the first device 100 and / or subtitle content obtained by performing speech recognition on audio content of the first device. The audio content may be participants' speech content during the communication, or may also be an audio file referred to by the shared video material; the screen content may be shared document content, a video file, PPT content, etc.
[0159] The translation module 1902 may translate, when the first device 100 has a translation capability, the target multimedia data to obtain translation content, and transmit the translation content to the second device 200 for display on the second device 200. For example, the translation module 1902 may perform translation continuity according to the AI capabilities of the first device 100 and the second device 200 using the first VCC app 110, including translation continuity of the audio content and / or translation continuity of the screen content.
[0160] According to an embodiment of the present disclosure, the translation module 1902 may process the target multimedia data, which may be audio content. The translation module 1902 may perform speech recognition on the audio content, for example, by using ASR technology, to obtain subtitle content, i.e., text content, corresponding to the audio content. The translation module 1902 may then translate the text content into text content in the target language.
[0161] In this way, by translating the audio content generated during the conference into text content in a specified language, it is convenient for a user at the second device 200 who does not speak the same language to promptly understand the specific content of the communication between the users participating in the conference, thereby further ensuring the smooth progress of the conference and improving the efficiency of the communication.
[0162] According to an embodiment of the present disclosure, the device 1900 for multimedia data transmission may further include a query module and a main device determining module.
[0163] The query module may query whether the first device 100 and / or the second device 200 has the translation capability. Then, the main device determining module may determine a main device for translating the target multimedia data according to a query result.
[0164] According to an embodiment of the present disclosure, the query module may transmit a translation capability query request to the second device 200. Then, the query module may receive information related to the translation capability of the second device 200 from the second device 200.
[0165] According to the embodiment of the present disclosure, when the first device 100 has the translation capability, the translation module 1902 may translate the target multimedia data to obtain translation content.
[0166] For the audio content part, the main device determining module may determine whether the first device 100 or the second device 200 is used to perform translation continuity of the audio content based on whether the first device 100 or the second device 200 supports the ASR capability. For example, when the first device 100 and the second device 200 both have the ASR capability, the main device determining module may determine that the first device 100 may be preferably used to perform ASR and translation. When only the first device 100 has the ASR capability and the second device 200 does not have the ASR capability, the main device determining module may determine that the first device 100 may be used to perform ASR and translation.
[0167] For the screen content part, before performing the translation continuity on the screen content, the main device determining module may transmit an instruction to the second device 200 to query whether the second device 200 has an "OCR and translation" capability. Then, the main device determining module may determine whether the first device 100 or the second device 200 is used to perform screen translation continuity based on whether the first device 100 or the second device 200 supports the "OCR and translation" capability. For example, when the first device 100 and the second device 200 both support the OCR capability and the translation capability, or when only the first device 100 has the OCR capability and the translation capability while the second device 200 does not have the OCR capability and the translation capability, the main device determining module may use, call, or invoke the "OCR and translation" capability of the first device 100 to perform character recognition and translation on the screen content currently displayed on the first device 100, and may transmit the translated content to the second device 200 for display through the second VCC app 210.
[0168] According to an embodiment of the present disclosure, the device 1900 for multimedia data transmission may further include a first multimedia data transmitting module.
[0169] The first multimedia data transmitting module may transmit, when the first device 100 does not have the translation capability and the second device 200 has the translation capability, the target multimedia data and a translation instruction to the second device 200, wherein the translation instruction is used to instruct the second device 200 to translate the target multimedia data.
[0170] For the audio content part, when the first device 100 does not have the ASR capability and only the second device 200 has the ASR capability, the first multimedia data transmitting module may transmit the audio content obtained from the first device 100 to the second device 200 through the first VCC app 110 for performing ASR and translation to further obtain original text content corresponding to the audio content and translated text content in a target (or specified) language.
[0171] For the screen content part, when the first device 100 does not support the OCR and translation capability and only the second device 200 supports the OCR and translation capability, the first multimedia data transmitting module may use or invoke the OCR and translation capability of the second device 200 to perform character recognition and translation on a video frame transmitted by the first device 100.
[0172] According to an embodiment of the present disclosure, the device 1900 for multimedia data transmission may further include a second multimedia data transmitting module.
[0173] The second multimedia data transmitting module may transmit, when the first device 100 and the second device 200 do not have the translation capability, the target multimedia data and a translation instruction to a third device 300 having the translation capability. The translation instruction is used to instruct the third device 300 to translate the target multimedia data, and the obtained translation content is transmitted to the second device 200. The third device 300 may be, but is not limited to, a server. For example, translated content may first be transmitted by the third device 300 to the second device 200; or the translated content may also firstly be transmitted by the third device 300 to the first device 100, and then transmitted by the first device 100 to the second device 200. The present disclosure does not make any specific definition thereon.
[0174] In this way, when the first device 100 and the second device 200 do not have the translation capability, the third device 300 having the translation capability may perform translation continuity, thereby avoiding a situation in which the communication is forced to be interrupted because both the first device 100 and the second device 200 do not have the translation capability, thus ensuring the smooth progress of the communication and improving its efficiency.
[0175] According to an embodiment of the present disclosure, the translation module 1902 may project the translation content to the second device 200. For example, the second device 200 may not have the AI translation capability or may not have the second VCC app 210 installed, and at this time, the target multimedia data needs to be translated using the AI translation capability of the first device 100, and the first device 100 may perform content transfer with the second device 200 by using a projection protocol. For example, the projection protocol may be, but is not limited to, a Miracast protocol.
[0176] According to an embodiment of the present disclosure, the device 1900 for multimedia data transmission may further include a first original content display module. The first original content display module may display original content on the first device 100. The original content may include screen content and / or subtitle content, and translation content corresponding to the original content may be used for display on the second device 200.
[0177] According to an embodiment of the present disclosure, the device 1900 for multimedia data transmission may further include a second original content display module and a first original content transmitting module.
[0178] The second original content display module may display original content on the first device 100. The original content may include the screen content and / or the subtitle content. The first original content transmitting module may transmit the original content to the second device 200, wherein the original content and translation content corresponding to the original content may be used for display on the second device 200.
[0179] According to an embodiment of the present disclosure, when the original content includes the subtitle content, the second original content display module may not display the subtitle content on the first device 100, and the first original content transmitting module may transmit the subtitle content and the translation content corresponding to the subtitle content to the second device 200, wherein the subtitle content and the translation content corresponding to the subtitle content may be used for display on the second device 200.
[0180] According to an embodiment of the present disclosure, the device 1900 for multimedia data transmission may further include an original content and translation content display module and a second original content transmitting module.
[0181] The original content and translation content display module may display original content and translation content corresponding to the original content on the first device 100. The original content may include the screen content and / or the subtitle content. The second original content transmitting module may transmit the original content to the second device 200. The original content and the translation content corresponding to the original content may be used for display on the second device 200.
[0182] The device 1900 for multimedia data transmission may be implemented as part of the first device 100 or as a separate device communicatively connected to the first device 100. In the case where the device 1900 is implemented within the first device 100, the multimedia data obtaining module 1901 and the translation module 1902 may be executed by the processor of the first device 100. Not only the multimedia data obtaining module 1901 and the translation module 1902, but also any other modules that may be included in the device 1900 may be executed by the processor of the first device 100. Alternatively, the device 1900 may be an external device or server configured to perform multimedia data obtaining and translation in cooperation with the first device 100.
[0183] FIG. 20 is a block diagram illustrating a device for multimedia data transmission according to an embodiment of the present disclosure.
[0184] Referring to FIG. 20, a device 2000 for multimedia data transmission may include a first translation content receiving module 2001 and a translation content display module 2002.
[0185] The first translation content receiving module 2001 may receive, when the first device 100 has the translation capability, translation content for target multimedia data from the first device 100 from the first device 100, wherein the translation content is obtained by translating the target multimedia data by the first device 100 in response to receiving an instruction for enabling multimedia data translation. The translation content display module 2002 may display the received translation content on a screen of the device 2000.
[0186] According to an embodiment of the present disclosure, the target multimedia data may include screen content of the first device 100 and / or subtitle content obtained by performing speech recognition on audio content of the first device 100. The audio content may be participants' speech content during the communication or an audio file referred to by the shared video material, and the screen content may be shared document content, a video file, PPT content, etc.
[0187] According to an embodiment of the present disclosure, the device 2000 for multimedia data transmission may further include a query request receiving module and a translation capability information transmitting module.
[0188] The query request receiving module may receive a translation capability query request from the first device 100. Then, the translation capability information transmitting module may, in response to or based on the translation capability query request, transmit information related to the translation capability of the second device 200 to the first device 100.
[0189] According to an embodiment of the present disclosure, the device 2000 for multimedia data transmission may further include a multimedia data receiving module and a translation module.
[0190] The multimedia data receiving module may receive, when the first device 100 does not have the translation capability and the second device 200 has the translation capability, the target multimedia data and a translation instruction from the first device 100. Then, in response to or based on the received translation instruction, the translation module may translate the target multimedia data to obtain the translation content.
[0191] According to an embodiment of the present disclosure, the device 2000 for multimedia data transmission may further include a second translation content receiving module.
[0192] The second translation content receiving module may receive, when the first device 100 and the second device 200 do not have the translation capability, the translation content for the target multimedia data, wherein the translation content is obtained by translating the target multimedia data by the third device 300 having a translation capability. The third device 300 may be, but is not limited to, a server. For example, translated content may be transmitted by the third device 300 to the second device 200; or the translated content may first be transmitted by the third device 300 to the first device 100, and then transmitted by the first device 100 to the second device 200. The present disclosure does not make any specific definition thereon.
[0193] In this way, when the first device 100 and the second device 200 do not have the translation capability, the third device 300 having the translation capability may perform translation continuity, thereby avoiding a situation in which the communication is forced to be interrupted because the first device 100 and the second device 200 both do not have the translation capability, thus ensuring the smooth progress of the communication and improve the efficiency of the communication.
[0194] According to an embodiment of the present disclosure, the first translation content receiving module 2001 may receive translation content projected from the first device 100. For example, the second device 200 may not have the AI translation capability or may not have the second VCC app 210 installed, and at this time, the target multimedia data needs to be translated by using the AI translation capability of the first device 100. The first device 100 may transfer the translated content to the second device 200 by using a projection protocol. For example, the projection protocol may be, but is not limited to, a Miracast protocol.
[0195] The translation content display module 2002 may display the translation content.
[0196] According to an embodiment of the present disclosure, the translation content display module 2002 may display the translation content corresponding to the original content, wherein the original content may include screen content and / or subtitle content. The original content may further be displayed on the first device 100.
[0197] According to an embodiment of the present disclosure, the device 2000 for multimedia data transmission may further include a first original content receiving module.
[0198] The first original content receiving module may receive original content from the first device 100, wherein the original content may include screen content and / or subtitle content, and the original content may also be displayed on the first device 100. In addition, the translation content display module 2002 may display both the original content and the translation content corresponding to the original content on the second device 200.
[0199] According to an embodiment of the present disclosure, the translation content display module 2002 may display, when the original content includes the subtitle content, the subtitle content and the translation content corresponding to the subtitle content on the second device 200, wherein the subtitle content may not be displayed on the first device 100.
[0200] According to an embodiment of the present disclosure, the device 2000 for multimedia data transmission may further include a second original content receiving module.
[0201] When original content and translation content corresponding to the original content are displayed on the first device 100, wherein the original content may include screen content and / or subtitle content, the second original content receiving module may receive the original content from the first device 100, and the translation content display module 2002 may display the original content and the translation content corresponding to the original content on the second device 200.
[0202] The device 2000 for multimedia data transmission shown in FIG. 20 may be implemented as part of the second device 200 or as a separate device communicatively connected to the second device 200. When implemented within the second device 200, the first translation content receiving module 2001 and the translation content display module 2002 may be executed by a processor of the second device 200. Not only the first translation content receiving module 2001 and the translation content display module 2002, but also any other modules that may be included in the device 2000 may be executed by the processor of the second device 200. Alternatively, the device 2000 may be an external device or server configured to cooperate with the second device 200 to receive and display translation content.
[0203] FIG. 21 is a block diagram illustrating an electronic device 2100 according to an embodiment of the present disclosure.
[0204] Referring to FIG. 21, the electronic device 2100 may include at least one memory 2101 and at least one processor 2102. The at least one memory 2101 stores instructions that, when executed by the at least one processor 2102, execute the method for multimedia data transmission according to the embodiment of the present disclosure.
[0205] The electronic apparatus 2100 may be a Personal Computer (PC) computer, a tablet device, a personal digital assistant, a smart phone, or any other device capable of executing the above instruction set. The electronic apparatus 2100 does not have to be a single electronic apparatus, and may also be any aggregate of devices or circuits that can execute the above-mentioned instructions (or instruction sets) individually or jointly. The electronic apparatus 2100 may also be a part of an integrated control system or a system manager, or may be a portable electronic device configured to be interconnected with local or remote device (e.g., via wireless transmission) via interfaces.
[0206] In the electronic apparatus 2100, the at least one processor 2102 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. For example, rather than limitation, the processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.
[0207] The at least one processor 2102 may be a single processing unit or several processing units, all of which may include multiple computing units. The at least one processor 2102 may be implemented as at least one processor (or at least one processor circuitry), one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and / or any electronic devices that manipulate signals based on operational instructions. Among other capabilities, the at least one processor 2102 may be configured to fetch and execute computer-readable instructions and data stored in the memory 2101.
[0208] The at least one processor 2102 may control all functions of the electronic device 2100. The at least one processor 2102 may perform the method for multimedia data transmission (e.g., the method shown in FIG. 4A, FIG. 4B, or the method performed by the first device 100 shown in FIG. 7, 8, 9, or 10) by executing instructions stored in the memory 2101. The at least one processor 2102 may control at least one configuration (e.g., a communication interface, I / O modules, and other modules) included in the electronic device 2100 by using instructions stored in the memory 2101 and / or data stored in the memory 2101. The at least one processor 2102 may perform the method for multimedia data transmission by controlling one or more modules and / or the at least one configuration using instructions stored in the memory 2101 and / or data stored in the memory 2101. The at least one processor 2102 may be configured to perform operations or functions corresponding to the method for multimedia data transmission performed by the first device 100 illustrated in FIG. 2.
[0209] The at least one processor 2102 may execute instructions or codes stored in the memory 2101, wherein the memory 2101 may also store data. Instructions and data may be transmitted and received over a network via a network interface device, wherein the network interface device may use any known transmission protocol.
[0210] The memory 2101 may be integrated with the processor 2102, for example, RAM or flash memory is arranged in an integrated circuit microprocessor or the like. In addition, the memory 2101 may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The memory 2101 and the processor 2102 may be operatively coupled, or may communicate with each other, for example, through an I / O port, a network connection or the like, so that the processor 2102 can read files stored in the memory.
[0211] The electronic apparatus 2100 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic apparatus 2100 may be connected to each other via a bus and / or a network.
[0212] The second device 200 may be configured as the electronic device 2100 shown in FIG. 21.
[0213] When the second device 200 is configured as the electronic device 2100 shown in FIG. 21, at least one processor 2102 included in the electronic device 2100 may perform the operations of the second device 200 illustrated in FIGS. 7, 8, 9, 10, or 18. In addition, it may also perform operations corresponding to those of the second device 200 shown in FIG. 2.
[0214] According to an embodiment of the present disclosure, there is further provided a computer readable storage medium, wherein instructions stored the computer readable storage medium, when executed by a processor of an electronic device, enable the electronic device to execute the above-described method for multimedia data transmission. Examples of the computer-readable storage medium include: Read Only Memory (ROM), Random Access Programmable Read Only Memory (PROM), Electrically Erasable Programmable Read Only Memory (EEPROM), Random Access Memory (RAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD- RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, blu-ray or optical disc storage, Hard Disk Drive (HDD), Solid State Drive (SSD), card storage (such as, multimedia cards, secure digital (SD) cards or extreme speed digital (XD) cards), magnetic tapes, floppy disks, magneto-optical data storage devices, optical data storage devices, hard disks, solid state disks, and any other devices that are configured to store computer programs and any associated data, data files and data structures in a non-transitory manner and provide the computer programs and any associated data, data files and data structures to the processor or computer so that the processor or computer may execute the computer programs. The instructions or computer programs in the above-mentioned computer-readable storage medium may run or execute in an environment deployed in a computer apparatus such as a client, a host, an agent device, a server, or the like. In addition, in one example, the computer program and any associated data, data files and data structures may be distributed on networked computer systems, so that the computer programs and any associated data, data files and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0215] According to an embodiment of the present disclosure, there may further be provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the method for multimedia data transmission according to the present disclosure.
[0216] A method and device for multimedia data transmission, an electronic device, a storage medium, and a computer program product according to the present disclosure may, in a multi-user communication scenario, where multiple users do not speak the same language as each other, enable an AI translation continuity function so that corresponding translation content may be displayed on a second device200 that requires translation while communicating through a first device 100. In this way, a user that requires translation may understand the specific content of the communication in time to further ensure smooth progress of the communication process, thereby improving the efficiency of the communication.
[0217] According to the embodiment of the present disclosure, by translating the audio content generated during the conference into text content in a target (or specified) language, it is convenient for a user at the second device 200 who does not speak the same language to timely understand the specific content of the communication between the users participating in the conference, thereby further ensuring the smooth progress of the conference process and improving the efficiency of the communication.
[0218] According to the embodiment of the present disclosure, when the first device 100 and the second device 200 do not have the translation capability, the third device 300 having the translation capability may further be enabled to perform translation continuity, to thereby avoid a situation when the communication is forced to be interrupted since the first device 100 and the second device 200 do not have the translation capability, that is, it is able to ensure the smooth progress of the communication and improve the efficiency of the communication.
[0219] According to the embodiment of the present disclosure, when the screen of the second device 200 is relatively smaller than the screen of the first device 100, an audio translation result may be displayed in a floating window positioned above the translation result of the screen content, thereby enabling efficient use of the screen space of the second device 200 and ensuring that the information display is both comprehensive and clear.
[0220] According to the embodiment of the present disclosure, when the screen of the second device 200 is relatively larger than the screen of the first device 100, a translation continuity result of the audio and a translation continuity result of the screen content may be displayed separately by means of screen division, thereby facilitating efficient use of the screen space of the second device 200 and ensuring that the information is displayed comprehensively and clearly.
[0221] According to the embodiment of the present disclosure, the determining whether the first device 100 is capable of translating the obtained target multimedia data comprises determining a device to perform translation from among the first device 100 and the at least one other device, based on translation capability of the first device 100 and the at least one other device.
[0222] According to the embodiment of the present disclosure, the translating the target multimedia data comprises: performing optical character recognition (OCR) on screen content to extract a first text; converting audio content into a second text using automatic speech recognition (ASR); translating the extracted first text from the screen content and the converted second text from the audio content into a target language; and generating translation display content by replacing the first text in the screen content with the translated first text, and generating subtitles corresponding to the translated second text, such that the translation display content and the subtitles are used for display on at least one other device.
[0223] According to the embodiment of the present disclosure, the method further may comprise querying whether the first device 100 and / or the at least one other device has the translation capability, and determining a main device for translating the target multimedia data according to a query result.
[0224] According to the embodiment of the present disclosure, the querying whether the at least one other device has the translation capability comprises: transmitting a translation capability query request to the at least one other device; and receiving information on the translation capability of the at least one other device from the at least one other device.
[0225] According to the embodiment of the present disclosure, the transmitting the translation content to the at least one other device may comprise projecting the translation content onto the at least one other device.
[0226] According to the embodiment of the present disclosure, the method may further comprise displaying original content on the first device 100, wherein the original content comprises the screen content and / or subtitle content obtained by performing speech recognition on audio content of the first device 100, and transmitting the original content to the at least one other device, wherein the original content and the translation content corresponding to the original content are used for display on the at least one other device.
[0227] According to the embodiment of the present disclosure, the processor 2102 may be further configured to, when determining whether the electronic device 2100 is capable of translating the obtained target multimedia data, determine a device to perform translation from among the electronic device 2100 and the at least one other device, based on translation capability of the electronic device 2100 and the at least one other device.
[0228] According to the embodiment of the present disclosure, the processor 2102 may be further configured to, when translating the target multimedia data, perform optical character recognition (OCR) on screen content to extract a first text, convert audio content into a second text using automatic speech recognition (ASR), translate the extracted first text from the screen content and the converted second text from the audio content into a target language, and generate translation display content by replacing the first text in the screen content with the translated first text, and generate subtitles corresponding to the translated second text, such that the translation display content and the subtitles are used for display on at least one other device.
[0229] According to the embodiment of the present disclosure, the processor 2102 may be further configured to query whether the electronic device and / or the at least one other device has the translation capability, and determine a host device for translating the target multimedia data according to a query result.
[0230] According to the embodiment of the present disclosure, the processor 2102 may be further configured to, when querying whether the at least one other device has the translation capability, transmit a translation capability query request to the at least one other device, and receive information on the translation capability of the at least one other device from the at least one other device.
[0231] According to the embodiment of the present disclosure, the processor 2102 may be further configured to, when transmitting the translation content to the at least one other device, project the translation content onto the at least one other device.
[0232] According to the embodiment of the present disclosure, the processor 2102 may be further configured to display original content on the electronic device, wherein the original content comprises the screen content and / or subtitle content obtained by performing speech recognition on audio content of the electronic device, and transmit the original content to the at least one other device, wherein the original content and the translation content corresponding to the original content are used for display on the at least one other device.
[0233] According to the embodiment of the present disclosure, the method may be further comprising: in a case where the first device 100 does not have the translation capability and the second device 200 has the translation capability, transmitting the target multimedia data and a translation instruction to the second device 200, wherein the translation instruction is used for instructing the second device 200 to translate the target multimedia data.
[0234] According to the embodiment of the present disclosure, the method may further comprise: in a case where the first device 100 and the second device 200 do not have the translation capability, transmitting the target multimedia data and a translation instruction to a third device 300 having a translation capability, wherein the translation instruction is used for instructing the third device 300 to translate the target multimedia data, and obtained translation content is transmitted to the second device 200.
[0235] According to the embodiment of the present disclosure, the method may further comprise: querying whether the first device 100 and / or the second device 200 has the translation capability; and determining a main device for translating the target multimedia data according to a query result.
[0236] According to the embodiment of the present disclosure, the querying whether the second device 200 has the translation capability may comprise transmitting a translation capability query request to the second device 200, and receiving information on the translation capability of the second device 200 from the second device 200.
[0237] According to the embodiment of the present disclosure, the transmitting the translation content to the second device 200 may comprise projecting the translation content onto the second device 200.
[0238] According to the embodiment of the present disclosure, the target multimedia data may comprise screen content of the first device 100 and / or subtitle content obtained by performing speech recognition on audio content of the first device 100.
[0239] According to the embodiment of the present disclosure, the method may further comprise displaying original content on the first device 100, wherein the original content comprises the screen content and / or the subtitle content, wherein the translation content corresponding to the original content is used for display on the second device 200.
[0240] According to the embodiment of the present disclosure, the method may further comprise displaying original content on the first device 100, wherein the original content comprises the screen content and / or the subtitle content, transmitting the original content to the second device 200, wherein the original content and the translation content corresponding to the original content are used for display on the second device 200.
[0241] According to the embodiment of the present disclosure, the method may further comprise displaying original content and the translation content corresponding to the original content on the first device 100, wherein the original content comprises the screen content and / or the subtitle content, and transmitting the original content to the second device 200, wherein the original content and the translation content corresponding to the original content are used for display on the second device 200.
[0242] According to the embodiment of the present disclosure, in a case where the original content comprises the subtitle content, the subtitle content is not displayed on the first device 100, and the subtitle content and the translation content corresponding to the subtitle content are used for display on the second device 200.
[0243] According to the embodiment of the present disclosure, the method may further comprise, in a case where the first device 100 does not have the translation capability and the second device 200 has the translation capability, receiving the target multimedia data and a translation instruction from the first device 100, and in response to the translation instruction, translating the target multimedia data to obtain the translation content.
[0244] According to the embodiment of the present disclosure, the method may further comprise, in a case where the first device 100 and the second device 200 both do not have the translation capability, receiving the translation content for the target multimedia data, wherein the translation content is obtained by translating the target multimedia data by a third device 300 having a translation capability.
[0245] According to the embodiment of the present disclosure, the method may further comprise receiving a translation capability query request from the first device 100, and in response to the translation capability query request, transmitting information on the translation capability of the second device 200 to the first device 100.
[0246] According to the embodiment of the present disclosure, the receiving the translation content for the target multimedia data in the first device 100 from the first device 100 may comprise receiving the translation content projected from the first device 100.
[0247] According to the embodiment of the present disclosure, the target multimedia data may comprise screen content of the first device 100 and / or subtitle content obtained by performing speech recognition on audio content of the first device 100.
[0248] According to the embodiment of the present disclosure, the displaying the translation content may comprise displaying the translation content corresponding to the original content on the second device 200, wherein the original content comprises the screen content and / or the subtitle content, wherein the original content is further displayed on the first device 100.
[0249] According to the embodiment of the present disclosure, the method may further comprise receiving the original content from the first device 100, wherein the original content comprises the screen content and / or the subtitle content, and the original content is further displayed on the first device 100, wherein the displaying the translation content may comprise displaying the original content and the translation content corresponding to the original content on the second device 200.
[0250] According to the embodiment of the present disclosure, the original content and the translation content corresponding to the original content are used for display on the first device 100, wherein the original content comprises the screen content and / or the subtitle content, the method may further comprise receiving the original content from the first device 100, the displaying the translation content may comprise displaying the original content and the translation content corresponding to the original content on the second device 200.
[0251] According to the embodiment of the present disclosure, in a case where the original content comprises the subtitle content, the displaying the original content and the translation content corresponding to the original content on the second device 200 may comprise display the subtitle content and the translation content corresponding to the subtitle content on the second device 200, the subtitle content is not displayed on the first device 100.
[0252] According to the embodiment of the present disclosure, the device 1900 may further include a first multimedia data transmitting module configured to, in a case where the first device 100 does not have the translation capability and the second device 200 has the translation capability, transmit the target multimedia data and a translation instruction to the second device 200, wherein the translation instruction is used for instructing the second device 200 to translate the target multimedia data.
[0253] According to the embodiment of the present disclosure, the device 1900 may further include a second multimedia data transmitting module configured to, in a case where the first device 100 and the second device 200 do not have the translation capability, transmit the target multimedia data and a translation instruction to a third device 300 having a translation capability, wherein the translation instruction is used for instructing the third device 300 to translate the target multimedia data, and obtained translation content is transmitted to the second terminal.
[0254] According to the embodiment of the present disclosure, the device 1900 may further include a query module configured to query whether the first device 100 and / or the second device 200 has the translation capability; and a main device determining module configured to determine a main device for translating the target multimedia data according to a query result.
[0255] According to the embodiment of the present disclosure, the query module is configured to transmit a translation capability query request to the second device 200; and receive information on the translation capability of the second device 200 from the second device 200.
[0256] According to the embodiment of the present disclosure, the translation module is configured to project the translation content onto the second device 200.
[0257] According to the embodiment of the present disclosure, the target multimedia data includes screen content of the first device 100 and / or subtitle content obtained by performing speech recognition on audio content of the first device 100.
[0258] According to the embodiment of the present disclosure, the device 1900 may further include a first original content display module configured to display original content on the first device 100, wherein the original content includes the screen content and / or the subtitle content, wherein the translation content corresponding to the original content is used for display on the second device 200.
[0259] According to the embodiment of the present disclosure, the device 1900 may further include a second original content display module configured to display original content on the first device 100, wherein the original content includes the screen content and / or the subtitle content, and a first original content transmission content configured to transmit the original content to the second device 200, wherein the original content and the translation content corresponding to the original content are used for display on the second device 200.
[0260] According to the embodiment of the present disclosure, the device 1900 may further include an original content and translation content display module configured to display original content and the translation content corresponding to the original content on the first device 100, wherein the original content includes the screen content and / or the subtitle content; and a second original content transmitting module configured to transmit the original content to the second device 200, wherein the original content and the translation content corresponding to the original content are used for display on the second device 200.
[0261] According to the embodiment of the present disclosure, in a case where the original content includes the subtitle content, the subtitle content is not displayed on the first device 100, and the subtitle content and the translation content corresponding to the subtitle content are used for display on the second device 200.
[0262] The technical solutions provided by the embodiment of the present disclosure at least bring the following advantageous effects: in the present disclosure, in a multi-user communication scenario, if multiple users do not speak the same language as each other, an AI translation continuity function may be enabled so that corresponding translation content can be displayed on the second device 200 that requires translation while communicating through the first device 100. In this way, a user that requires translation may understand specific content of communication in time to further ensure smooth progress of the communication process, thereby improving the efficiency of communication.
[0263] Those skilled in the art will easily conceive of other implementation solutions of the present disclosure after considering the description and practicing the invention disclosed herein. The present disclosure is intended to cover any modifications, uses, or adaptive changes of the present disclosure. These modifications, uses, or adaptive changes follow the general principles of the present disclosure and include common knowledge or customary technical means in the technical field that are not disclosed by the present disclosure. The description and the embodiment are only regarded as exemplary, and the true scope and spirit of the present disclosure are indicated by the claims below.
[0264] It should be understood that the present disclosure is not limited to the above-described accurate structures as illustrated in the accompanying drawings, and various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is only defined by the appending claims.
Claims
1.A method for multimedia data transmission performed by a first device 100, the method comprising:based on receiving an instruction for enabling multimedia data translation, obtaining target multimedia data by the first device 100;determining whether the first device 100 is capable of translating the obtained target multimedia data;translating the target multimedia data to obtain translation content when it is determined that the first device 100 is capable of translating the obtained target multimedia data, andtransmitting the translation content to at least one other device for displaying the translation content on the at least one other device.2.The method of claim 1, wherein the determining whether the first device 100 is capable of translating the obtained target multimedia data comprises:determining a device to perform translation from among the first device 100 and the at least one other device, based on translation capability of the first device 100 and the at least one other device.3.The method of claim 1 or 2, wherein the translating the target multimedia data comprises:performing optical character recognition (OCR) on screen content to extract a first text;converting audio content into a second text using automatic speech recognition (ASR);translating the extracted first text from the screen content and the converted second text from the audio content into a target language; andgenerating translation display content by replacing the first text in the screen content with the translated first text, and generating subtitles corresponding to the translated second text such that the translation display content and the subtitles are used for display on at least one other device.4.The method of any one of claims 1 to 3, further comprising:querying whether the first device 100 and / or the at least one other device has the translation capability; anddetermining a main device for translating the target multimedia data according to a query result.5.The method of claim 4, wherein the querying whether the at least one other device has the translation capability comprises:transmitting a translation capability query request to the at least one other device; andreceiving information on the translation capability of the at least one other device from the at least one other device.6.The method of any one of claim 1 to 5, wherein the transmitting the translation content to the at least one other device comprises:projecting the translation content onto the at least one other device.7.The method of any one of claims 1 to 6, further comprising:displaying original content on the first device 100, wherein the original content comprises the screen content and / or subtitle content obtained by performing speech recognition on audio content of the first device 100; andtransmitting the original content to the at least one other device, wherein the original content and the translation content corresponding to the original content are used for display on the at least one other device.8.An electronic device 2100 comprising:at least one processor 2102; anda memory 2101 for storing at least one instruction executable by the at least one processor 2102,wherein the at least one processor 2102 is configured to execute the at least one instruction to:based on receiving an instruction for enabling multimedia data translation, obtain target multimedia data;determine whether the electronic device 2100 is capable of translating the obtained target multimedia data;translate the target multimedia data to obtain translation content when it is determined that the electronic device 2100 is capable of translating the obtained target multimedia data; andtransmit the translation content to at least one other device for displaying the translation content on the at least one other device.9.The electronic device 2100 of claim 8, wherein the at least one processor 2102 is further configured to, when determining whether the electronic device 2100 is capable of translating the obtained target multimedia data, determine a device to perform translation from among the electronic device 2100 and the at least one other device, based on translation capability of the electronic device 2100 and the at least one other device.10.The electronic device 2100 of claim 8 or 9, wherein the at least one processor 2102 is further configured to, when translating the target multimedia data:perform optical character recognition (OCR) on screen content to extract a first text;convert audio content into a second text using automatic speech recognition (ASR);translate the extracted first text from the screen content and the converted second text from the audio content into a target language; andgenerate translation display content by replacing the first text in the screen content with the translated first text, and generate subtitles corresponding to the translated second text such that the translation display content and the subtitles are used for display on at least one other device.11.The electronic device 2100 of any one of claims 8 to 10, wherein the at least one processor 2102 is further configured to:query whether the electronic device 2100 and / or the at least one other device has the translation capability; anddetermine a main device for translating the target multimedia data according to a query result.12.The electronic device 2100 of claim 11, wherein the at least one processor 2102 is further configured to, when querying whether the at least one other device has the translation capability:transmit a translation capability query request to the at least one other device; andreceive information on the translation capability of the at least one other device from the at least one other device.13.The electronic device 2100 of any one of claim 8 to 12, wherein the at least one processor 2102 is further configured to, when transmitting the translation content to the at least one other device, project the translation content onto the at least one other device.14.The electronic device 2100 of any one of claims 8 to 13, wherein the at least one processor 2102 is further configured to:display original content on the electronic device 2100, wherein the original content comprises the screen content and / or subtitle content obtained by performing speech recognition on audio content of the electronic device 2100; andtransmit the original content to the at least one other device, wherein the original content and the translation content corresponding to the original content are used for display on the at least one other device.15.A computer readable storage medium storing instructions that, when executed by at least one processor 2102, cause the at least one processor 2102 to perform the method of any one of claims 1-7.