Electronic device and method for multi-speaker call
The electronic device addresses the challenge of language barriers in multi-party calls by translating speech to text, identifying and explaining unfamiliar entities, thereby improving communication clarity.
Patent Information
- Application Number
- PCT/KR2025/000507
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-17
AI Technical Summary
Existing multi-party calling systems struggle with accurate real-time translation and providing additional explanations for unfamiliar words or entities during conversations among participants speaking different languages.
An electronic device equipped with a communication circuit, memory, and processor that translates speech information into text, identifies entities requiring additional explanation, obtains description information, and transmits the translated text and additional descriptions to external devices for multi-party calls.
Facilitates real-time translation and provides additional explanations for unknown entities, enhancing understanding among multi-party call participants by ensuring accurate and comprehensive communication.
Smart Images

Figure KR2025000507_17072025_PF_FP_ABST
Abstract
Description
Electronic device and method for multi-party communication
[0001] The present disclosure relates to an electronic device and method for multi-party calling.
[0002] With the advancement of digital technology, electronic devices are now available in various forms, such as smartphones, tablet personal computers (PCs), and personal digital assistants (PDAs). Electronic devices are also being developed into wearable devices to enhance portability and accessibility.
[0003] Electronic devices are being developed to enable real-time multi-speaker conversations between people of different nationalities for various purposes, such as meetings, conferences, and games. In multi-speaker conversations, participants can either speak in a unified English language or speak in their own language, with real-time translation provided as subtitles in their own language. For example, if a participant speaking a certain language includes recently popular or newly coined Korean words, the translation may not be accurate, and even if translated, it may be difficult for participants speaking other languages to understand.
[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0005] According to one embodiment of the present disclosure, an electronic device includes a communication circuit, a memory storing instructions, and at least one processor.
[0006] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to obtain a first text converted from speech information in a first language.
[0007] According to one embodiment, the instructions, when executed by the at least one processor, may be configured to cause the electronic device to generate a second text that is a translation of the first text into a second language.
[0008] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to identify a first entity in the first text that requires additional description in the second language.
[0009] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to obtain first description information describing the first entity name in the first language.
[0010] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to generate second description information in the second language for the additional description based on the first description information.
[0011] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to control the communication circuit to transmit the second text and the second description information to an external electronic device.
[0012] According to one embodiment, a method of operating in an electronic device includes obtaining a first text converted from speech information in a first language.
[0013] According to one embodiment, a method of operation in an electronic device includes generating a second text translated from the first text into a second language.
[0014] According to one embodiment, a method of operation in an electronic device may include identifying a first entity in the first text that requires additional description in the second language.
[0015] According to one embodiment, a method of operating in an electronic device includes obtaining first description information that describes the first entity name in the first language.
[0016] According to one embodiment, a method of operating in an electronic device includes generating second description information in a second language for the additional description based on the first description information.
[0017] According to one embodiment, a method of operating in an electronic device includes transmitting the second text and the second description information to an external electronic device.
[0018] According to one embodiment, an electronic device includes a display, a communication circuit, a memory storing instructions, and at least one processor.
[0019] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to receive speech information in a first language from the first electronic device, generate a second text by translating a first text in the first language into a second language by converting the received speech information, identify a first entity requiring additional description in the second language from the first text, obtain first description information that describes the first entity in the first language, generate second description information in the second language for the additional description based on the first description information, and control the display to display the second text in the second language and the second description information in the second language.
[0020] According to one embodiment, in a non-transitory storage medium storing a program, the program includes executable instructions that, when executed by at least one processor of an electronic device, cause the electronic device to perform an operation of obtaining a first text converted from speech information in a first language, an operation of generating a second text translated from the first text into a second language, an operation of identifying a first entity requiring additional description in the second language from the first text, an operation of obtaining first description information describing the first entity in the first language, an operation of generating second description information in the second language for the additional description based on the first description information, and an operation of transmitting the second text and the second description information to an external electronic device.
[0021] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0022] FIG. 2 is a drawing showing an example configuration of an electronic device according to one embodiment.
[0023] Figure 3 is a diagram showing an example of a configuration of a conversation module according to one embodiment.
[0024] Figure 4 is a drawing showing an example of operation in a conversation module according to one embodiment.
[0025] FIG. 5 is a diagram showing an example of operation in a conversation module according to one embodiment.
[0026] Figure 6 is a drawing showing an example of operation in a conversation module according to one embodiment.
[0027] Figure 7 is a drawing showing an example of operation in a conversation module according to one embodiment.
[0028] Figure 8 is a drawing showing an example of operation in a conversation module according to one embodiment.
[0029] FIG. 9 is a diagram showing an example of operation in a conversation module according to one embodiment.
[0030] FIG. 10A and FIG. 10B are diagrams showing examples of operations in a conversation module according to one embodiment.
[0031] FIG. 11 is a diagram showing an example configuration of an electronic device and server for multi-party calling according to various embodiments.
[0032] FIG. 12 is a drawing showing an example of an operating method in an electronic device according to one embodiment.
[0033] FIG. 13 is a diagram illustrating an example for multi-party calling in an electronic device according to one embodiment.
[0034] FIG. 14 is a diagram illustrating an example for multi-party calling in an electronic device according to one embodiment.
[0035] FIG. 15 is a drawing showing an example of an operating method in an electronic device according to one embodiment.
[0036] FIG. 16 is a diagram illustrating an example of a multi-party call in an electronic device according to one embodiment.
[0037] FIG. 17 is a diagram illustrating an example of a multi-party call in an electronic device according to one embodiment.
[0038] FIG. 18 is a diagram illustrating an example of an operation method between electronic devices and a server according to one embodiment.
[0039] FIG. 19 is a drawing showing an example of an operating method in a second electronic device according to one embodiment.
[0040] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0041] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components. In addition, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness. The term "user" used in the embodiments of the present disclosure may refer to a person using an electronic device or a device (e.g., an artificial intelligence electronic device) using an electronic device.
[0042] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0043] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0044] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0045] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0046] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0047] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0048] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0049] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0050] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0051] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0052] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0053] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0054] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0055] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0056] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0057] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0058] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0059] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for realizing 1eMBB, a loss coverage (e.g., 164 dB or less) for realizing mMTC, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for realizing URLLC.
[0060] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0061] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0062] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0063] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0064] FIG. 2 is a diagram illustrating a configuration of an electronic device according to one embodiment.
[0065] Referring to FIG. 2, an electronic device (201) according to one embodiment (e.g., the electronic device (101) of FIG. 1) may include at least one processor (210), memory (220), communication circuit (230), audio circuit (240), microphone (250), display (260), and speaker (270). Without being limited thereto, the electronic device (201) may further include other components described in FIG. 1.
[0066] According to one embodiment, at least one processor (210) (hereinafter referred to as a processor) of the electronic device (201) can control the overall operation of the electronic device (201) and can be implemented as one or more processors. The processor (220) can be implemented in the same or similar manner as the processor (120) of FIG. 1.
[0067] According to one embodiment, the processor (210) can control a conversation module (e.g., a device, a component, or a circuit) for a multi-talker call (e.g., a video call or a conversation) installed in the electronic device (201) or stored in the memory (220), and, if at least one word included in voice information spoken during the multi-talker call is a word whose meaning is unknown in another language, the processor (210) can perform an operation to provide an additional explanation for the word through the conversation module (301).
[0068] According to one embodiment, the processor (210) may control at least one other component (e.g., a hardware or software component) of the electronic device (201) connected to the processor (210) by executing software, for example, and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (210) may store commands or data received from other components (e.g., a display (260) or a communication circuit (230)) in the memory (220), process the commands or data stored in the memory (220), and store result data in the memory (220).
[0069] According to one embodiment, the processor (210) may obtain voice information (e.g., a voice signal) spoken by a user (e.g., a speaker) received through a microphone (250) (e.g., an input module (150) of FIG. 1).
[0070] According to one embodiment, the processor (210) can generate a second text translated from a first text in a first language into a second language.
[0071] According to one embodiment, the processor (210) may identify a first named entity (e.g., an entity) in a first text that requires additional description in a second language (not detected in the second language). According to one embodiment, the processor (210) may identify (e.g., identify or extract) at least one named entity in the first language corresponding to at least one word included in the first text, and may identify (e.g., select or identify) a first named entity requiring additional description among the identified at least one named entity in the first language. According to one embodiment, the processor (210) may identify (e.g., identify or extract) at least one named entity in the first language from a first named entity database using a first method (e.g., a rule-based method), and may identify (e.g., identify or extract) at least one named entity in the second language for at least one word included in the second text from a second named entity database, and may compare the identified at least one named entity (e.g., a translated named entity) with the identified at least one named entity in the second language. The processor (210) may select a first language entity name that does not match at least one second language entity name among at least one first language entity name. Here, the selected first language entity name may be an entity name that is included in the first entity database, and whose translated entity name is not searched in the second entity database. According to one embodiment, the processor (210) may use a second method (e.g., a deep-based or deep learning based algorithm) to check the similarity between at least one first language entity name (or entity information) and at least one second language entity name (or entity information), and may select a first language entity name whose confirmed similarity value is greater than or equal to a pre-specified threshold value as an entity name that does not require additional explanation. The processor (210) may select a first language entity name whose confirmed similarity value is less than or equal to a pre-specified threshold value as an entity name that requires additional explanation.According to one embodiment, the processor (210) can identify the entity name of the first language selected using the first method and selected using the second method as the first entity name.
[0072] In one embodiment, the processor (210) may obtain first description information in a first language for additional description of the first entity name in the identified first language. The processor (210) may obtain (e.g., determine or generate) data of the identified first entity name from the first entity name database as the first description information.
[0073] According to one embodiment, the processor (210) may generate (e.g., obtain) second description information in a second language based on the first description information. The processor (210) may translate the first description information into the second language and refine the translated description information to generate the second description information in the second language.
[0074] According to one embodiment, the processor (210) may control the communication circuit (230) to transmit result information including second text (e.g., a translation of the first text) and second descriptive information in a second language to an external electronic device.
[0075] According to one embodiment, the memory (220) may store information related to the operation of the electronic device (201), information for multi-talker calls between external electronic devices (e.g., real-time video information about connected users, an execution screen for multi-talker calls, information related to functions provided on the execution screen), voice information input through a microphone (250), first text converted from the voice information, a selected first entity name (e.g., first NE), and first description information. According to one embodiment, the memory (220) may store a first entity name database (e.g., first NE DB) preset to provide additional description. The memory (220) may store a second entity name database (e.g., a second NE DB) received from a server (e.g., a server (108) of FIG. 1) or an external electronic device (e.g., an electronic device (104) of FIG. 1). Without being limited thereto, according to one embodiment, the memory (220) may not store the first entity name database and / or the second entity name database, but may communicate with a server (e.g., a server (108) of FIG. 1) that stores the first entity name database and / or the second entity name database to store only result information (e.g., NE verification result information) according to a request. According to one embodiment, the memory (220) may store translated text (e.g., a text translated from voice information according to a user's speech of the external electronic device) and additional description (e.g., description information of an entity detected as not being verified in a first language set in the electronic device) received from the external electronic device.
[0076] According to one embodiment, the communication circuit (230) (e.g., the communication module (190) of FIG. 1) may be configured to perform a communication function between the electronic device (201) and the external electronic device (101). The communication circuit (230) may transmit information related to a multi-talker call and voice information input according to the speaker's speech in the electronic device (201) using a wireless communication method, and may transmit an additional description of an entity name of a word not identified in another language among at least one word included in a text translated from the voice information.
[0077] According to one embodiment, the audio circuit (240) (e.g., the audio module (170) of FIG. 1) may, under the control of the processor (210), process voice information (or voice signal) input from the microphone (250) and voice information (or voice signal) received from an external electronic device and output a processed output signal for the user to hear through the speaker (270).
[0078] According to one embodiment, the display (260) (e.g., the display module (160) of FIG. 1) may display an execution screen of an application related to a multi-party call. The display (260) may, under the control of the processor (210), display result information including information related to the multi-party call, user screens of users participating in the multi-party call, and translations and additional descriptions of spoken voice information. The display (260) may, under the control of the processor (210), display result information including information related to users, translations and additional descriptions of user screens and spoken voice information when performing a multi-party call. According to one embodiment, the display (260) may be implemented in the form of a touch screen. When the display (260) is implemented together with an input module in the form of a touch screen, it may display various pieces of information generated according to a user's touch operation. According to one embodiment, the display (260) may be configured with at least one of a liquid crystal display (LCD), a thin film transistor LCD (TFT-LCD), organic light emitting diodes (OLED), a light emitting diode (LED), an active matrix organic LED (AMOLED), a flexible display, and a 3-dimensional display. In addition, some of these displays may be configured as transparent or light-transmitting so that the outside can be seen through them. This may be configured in the form of a transparent display including a transparent OLED (TOLED). According to another embodiment, in addition to the display (260), other display modules (e.g., an extended display or a flexible display) may be further installed.
[0079] As such, in one embodiment, the main components of the electronic device have been described through the electronic device (101) of FIG. 1 and the electronic device (201) of FIG. 2. However, in various embodiments, not all of the components illustrated through FIGS. 1 and 2 are essential components, and the electronic device (101, 201) may be implemented with more components than the illustrated components, or may be implemented with fewer components. In addition, the positions of the main components of the electronic device (101, 201) described above through FIGS. 1 and 2 may be changed according to various embodiments.
[0080] FIG. 3 is a drawing showing an example of a configuration of a conversation module according to one embodiment, and FIGS. 4, 5, 6, 7, 8, 9, 10a, and 10b are drawings showing examples of operations in a conversation module according to one embodiment.
[0081] Referring to FIGS. 2, 3, 4, 5, 6, 7, 8, 9, 10a and 10b, according to one embodiment, a conversation module (e.g., a device, a component or a circuit) (301) may be stored in a memory (220) of an electronic device (201) and executed by a processor (210), or may be included in a server (e.g., the server (108) of FIG. 1) and executed by a processor of the server. For example, the conversation module (301) may be stored in a memory (220) of an external electronic device (e.g., a second electronic device).
[0082] According to one embodiment, the conversation module (301) may perform an operation to provide an additional explanation for a word included in speech information spoken during a multi-talk call (e.g., a video call or conversation) if the word is a word whose meaning is unknown in another language.
[0083] According to one embodiment, the conversation module (301) may include a voice recognition module (310), a context verification module (320), a language translation module (330), an entity name verification module (340), and a result information provision module (350).
[0084] According to one embodiment, the voice recognition module (310) may convert voice information spoken by a user (e.g., a speaker) into first text in a first language. The voice recognition module (310) may transmit the voice information in the first language spoken by the user (e.g., a speaker) to an external electronic device (e.g., a second electronic device). Here, the voice information may be transmitted in synchronization with the result information. The present invention is not limited thereto, and the conversation module (301) may not include the voice recognition module (310).
[0085] According to one embodiment, the context verification module (320) can obtain the context of the current first text based on the conversation history accumulated in the history database (DB) (361). The context verification module (320) can obtain the context of the accumulated conversation content from among the classified context models. Here, the classified context modules can be configured by classifying the text using a context classification model (e.g., a model with a structure such as BERT (bidirectional encoder representations from transformers) or ELECTRA (efficiently learning an encoder that classifies token replacements accurately)) (363). The context verification module (320) sets the window size to identify only the context of the recent conversation, since it is difficult to input the conversation history into the model when it increases, and there is a possibility that the context of the conversation may change when the conversation is long. Thus, only the most recent conversation can be identified and provided. The context verification module (320) can determine sentence context by analyzing conversation content belonging to a window of a specified length (e.g., 100 tokens, 5 sentences), and the window can move at specified intervals (e.g., 20 tokens, 1 sentence) to track changes in the sentence context.
[0086] According to one embodiment, the context verification module (320) may generate, as output data, a context vector value (e.g., [AI, semiconductor, …] = [0.9, 0.5, …]) composed of the probability of each context based on the accumulated conversation content (e.g., “Among the papers recently published by MIT, there is a technology developed to enable AI to learn faster. It was developed in collaboration with NVIDIA and the research team announced that it is planned to be installed on the Google platform”) using a history database (361), as illustrated in FIG. 4. Here, the context vector value may be generated using at least one value that satisfies a threshold value or more among the probabilities of each context. For example, the conversation module (301) may learn and store models for each context using a designated model (e.g., multi-model), as illustrated in FIG. 5, and select a model that fits the context during inference. For example, the conversation module (301) can store a model trained to recognize various classes in a single model using another designated model (e.g., multi-value one model). Here, the other designated model (e.g., multi-value one model) can be placed at the beginning of the conversation content during learning / inference, such as a multilingual translation model. <context>" can be added, and it is one model, but when translating, it is at the very front <context>The model can decide on its own how to translate the following sentence by referring to the part. For example, the context verification module (320) can check the context (e.g., input) at the very beginning of the input speech information (e.g., input). <ai>) can be added.
[0087] According to one embodiment, the language translation module (330) acquires the context (e.g., <ai>) can be used to translate a first text in a first language into a second text in a second language. The language translation module (330) can more accurately process homonyms that fit the conversation context by using the acquired context. The language translation module (330) can perform language translation using a translation model (365). For example, the language translation module (330) can use a transformer method, which is a text-to-text model (both input and output are text) consisting of an encoder that understands an input text and a decoder that generates a result (e.g., a translation text). Without being limited thereto, a transformer model with a speech-to-text or speech-to-speech structure can be used.
[0088] According to one embodiment, the named entity verification module (340) can verify (e.g., extract or identify) at least one named entity (NE) (e.g., object or entity) for at least one word included in a first text of a first language, and select (e.g., extract, verify or identify) an NE requiring additional explanation among the at least one NE. The named entity verification module (340) can perform an operation to select an NE requiring additional explanation because there is no need to provide additional explanation for the NE for all words included in the first text. According to one embodiment, the named entity verification module (340) can select an NE requiring additional explanation using, for example, a rule-based method (601) and / or a deep-based method (603), as illustrated in FIG. 6. According to one embodiment, if an NE requiring additional explanation is not identified in the first text, the named entity verification module (340) may not provide the additional explanation.
[0089] According to one embodiment, the entity name verification module (340) can verify (e.g., extract, select, or identify) an entity (e.g., a first entity) requiring additional explanation by comparing at least one entity in a first language identified in a first text with at least one entity in a second language identified in a second text using a rule-based method (601).
[0090] According to one embodiment, the entity verification module (340) may obtain (e.g., extract) (610) at least one entity in a first language for at least one word identified in a first text from a first entity database (NE DB) (367a) of the first language among entity databases (NE DBs) (367), and may obtain (e.g., extract) (630) at least one entity in a second language for at least one word identified in a second text from a second entity database (367b) of the second language. According to one embodiment, the entity verification module (340) may translate (620) the obtained at least one entity in the first language, and compare the translated at least one entity in the first language with the obtained at least one entity in the second language to select (e.g., extract, verify, or identify) (640) the entity in the first language (e.g., the first entity) that does not match the at least one entity in the second language among the at least one entity in the first language. For example, the selected first language entity name (e.g., first entity name) may be an entity name that is included in the first entity database (367a) and is not included in the second entity database (367b), and may be an entity name that requires additional explanation. According to one embodiment, the entity name verification module (340) may verify the selected first language entity name (e.g., first entity name) as an entity name that requires additional explanation, and may obtain description information for additional explanation of the selected first language entity name (e.g., first entity name). For example, as illustrated in FIG. 7, the entity verification module (340) can verify that the word "IU" included in the first text matches or is similar to the entity "IU" or the translation of "IU" in the first language identified in the first entity database (367a) of the first language (e.g., Korean NE DB) and the entity "IU" in the second language identified in the second entity database (367b) of the second language (e.g., American NE DB).The entity verification module (340) can determine (e.g., determine) that additional explanation for the entity "IU" in the first language is not necessary based on whether the entity "IU" in the first language, "IU", or the translation of "IU" ("IU"), matches or is similar to the entity "IU" in the second language. For example, as illustrated in FIG. 7, the entity verification module (340) can determine (e.g., determine) that additional explanation for the entity "IU" in the first language, "Loossemble", or the translation of "Loossemble", "Loossemble", is necessary if, for the word "Loossemble" included in the first text, a matching or similar entity "Loossemble" in the first language, or the translation of "Loossemble", is not confirmed in the first entity database (e.g., Korean NE DB) of the first language. For example, when another external electronic device with a third language set participates in a video call or conversation, the entity name verification module (340) can verify, as shown in FIG. 7, that the entity name "IU" (or a translation of "IU") in the first language and the entity name "Russemble" (or a translation of "Russemble") in the first language are not included in the third language entity name database (367) (e.g., American NE DB) of the third language, and can verify that additional explanation is required.
[0091] According to one embodiment, the entity verification module (340) can verify sentence similarity using, for example, a deep-based method (e.g., a deep learning algorithm using a BERT embedding model) (603), as illustrated in FIG. 8, to verify whether at least one entity in the first language identified in the first text requires additional explanation (650, 660). The entity verification module (340) can compare the similarity between vector values for entity data in the first language (e.g., "title", "domain", "description") for the word "IU" included in the first text and vector values for entity data in the second language (e.g., "title", "domain", "description"). When the entity name verification module (340) confirms that a similarity comparison result value (e.g., 0.9115) is greater than a threshold value, it determines that the output second language entity name (e.g., "IU") corresponding to the similarity result value (e.g., 0.9115) is similar to the input word "IU", and that the word "IU" has a meaning that can be understood by the second language, and that no additional explanation is required. The entity name verification module (340) determines that the output second language entity name (e.g., "BTS") corresponding to the similarity result value (e.g., 0.1737) is not similar to the input word "IU", and that the word "IU" has an unknown meaning in the second language, and can determine that additional explanation is required, if only a similarity result value (e.g., 0.1737) that is less than the threshold value is confirmed. For example, the BERT embedding model may be a model that outputs a single vector as an output when a sentence is input, and can infer the most similar classification result value (e.g., similarity result value) using the vector.
[0092] According to one embodiment, the result information providing module (350) may transmit result information (e.g., a translation) including a second text obtained from a language translation module (330) translating a sentence in a first language spoken by a user into a second language and second description information (e.g., an NE translation) in a second language for a selected entity (e.g., a first entity name) in the first language obtained from an entity name verification module (340) to an external electronic device. Here, the second text may be retrieved from a DB stored in the memory (220) or may be generated using a machine translator. The result information providing module (350) may store the result information in a history database (361).
[0093] According to one embodiment, the result information providing module (350) can translate data (or original entity name) of a selected first language entity name (e.g., first entity name) (e.g., "title":"Loossemble", "domain":"singer", "description":"It is a five-member girl group active in Korea") into a second language, and generate second description information in the second language based on a refined version (e.g., "Loossemble: It is a five-member girl group active in Korea") of the translated data. In the data of the entity name (e.g., the first entity name) of the first language, the domain and description can obtain the translation result using the translation model. For example, in the case of the title "Loossemble", it may not be translated into "Loossemble" using the translation model. Therefore, when pre-storing the data of the entity name (e.g., the first entity name) of the first language in the first entity name database (367a), for example, translation information (e.g., "en-us": "Loossemble" and / or "jp": "ルセンブル") in at least one other language (e.g., a language of another region) for the title "Loossemble" can be stored. For example, the translation information in at least one other language (e.g., a language of another region) is a result that comes out when "Loossemble" is directly translated, and can obtain the same effect as translating the phonetic notation into another language. According to one embodiment, the result information providing module (350) can transmit result information including a second text translated from a first text in a first language (e.g., "I think IU and Loossemble songs are good these days. *Loossemble: It is a five-member girl group active in Korea.") and descriptive information in the second language to an external electronic device, as illustrated in FIG. 10b.For example, since "Loossemble" is a singer who debuted not long ago and its meaning is unknown in a second language, the second language descriptive information for "Loossemble" (e.g., *Loossemble: It is a five-member girl group active in Korea.") can be added to the result information. For example, since "IU" is a singer who debuted long ago and has been active in a region of the second language, its meaning is unknown in a second language when translated as "IU," the second language descriptive information for "IU" can be omitted.
[0094] The operations in the dialogue module (301) described in the above-described drawings 3 to 10b (e.g., an interpretation dialogue module included in a software module) may be equally executed by the processor (210) of the electronic device (201) or by the dialogue module (301) installed in the electronic device (201) or stored in the memory (220) (e.g., included in a software module) under the control of the processor (210).
[0095] An electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2) according to one embodiment may implement a software module (e.g., the program (140) of FIG. 1) for a multi-party call (e.g., a video call or a conversation). A memory (e.g., the memory (130) of FIG. 1, the memory (220) of FIG. 2) of the electronic device may store commands (e.g., instructions) to implement the software module. At least one processor (e.g., the processor (120) of FIG. 1, the processor (210) of FIG. 2) may execute the commands stored in the memory to implement the software module and control hardware (e.g., the display module (160), the audio module (170), or the communication module (190) of FIG. 1) associated with a function of the software module.
[0096] A software module of an electronic device according to one embodiment may be configured to include a kernel (or a hardware abstraction layer (HAL)), a framework (e.g., middleware (144) of FIG. 1), and an application (e.g., application (146) of FIG. 1). At least a portion of the software module may be preloaded on the electronic device or may be downloadable from a server (e.g., server (108)).
[0097] In one embodiment, the kernel may include, but is not limited to, a system resource manager or device driver related to multi-party calling (e.g., video calling or chatting), and may further include other modules. The system resource manager may perform control, allocation, or retrieval of system resources. The device driver may include, for example, a display driver, a camera driver, a Bluetooth driver, a shared memory driver, a USB driver, a keypad driver, a WIFI driver, an audio driver, or an inter-process communication (IPC) driver.
[0098] According to one embodiment, the framework may be configured to include, for example, a conversation module (301), but is not limited thereto, and may further include other modules. The framework may provide functions commonly required by applications or provide various functions to applications through an application programming interface (API) (not shown) so that applications can efficiently use limited system resources within an electronic device. The framework may include modules that form a combination of various functions of components. The framework may provide specialized modules for each type of operating system to provide differentiated functions. The framework may dynamically delete some existing components or add new components.
[0099] According to one embodiment, the application may be configured to include applications (e.g., modules, managers, or programs) related to multi-party calls. For example, the application may include a module (or application) (not shown) for providing additional explanations for words (e.g., slang and / or neologisms) that do not exist in other languages during multi-party calls. The application may be configured to include a module (or application) (not shown) for communicating with an external electronic device (e.g., the electronic device (102, 104) or the server (108) of FIG. 1). The application may include applications received from the external electronic device (e.g., the server (108) or the electronic device (102 or 104)). According to one embodiment, the application may include a preloaded application or a third-party application downloadable from the server. The components and names of the components of the software module according to the illustrated embodiment may vary depending on the type of operating system. According to one embodiment, at least a portion of the software module may be implemented as software, firmware, hardware, or a combination of at least two or more thereof. At least a portion of the software module may be implemented (e.g., executed) by, for example, a processor (e.g., an AP). At least a portion of the software module may include, for example, at least one of a module, a program, a routine, a set of instructions, or a process for performing at least one function.
[0100] FIG. 11 is a diagram showing an example configuration of an electronic device and server for multi-party calling according to various embodiments.
[0101] Referring to FIGS. 3 and 11, an electronic device (201) (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIGS. 2 and 11) (hereinafter referred to as a first electronic device) and an external electronic device (1101) (e.g., the electronic device (104) of FIG. 1) (hereinafter referred to as a second electronic device) according to one embodiment are connected to a server (1103) (e.g., the server (108) of FIG. 1) through a communication function and can perform a multi-party call (e.g., a video call or conversation). The second electronic device (1101) (e.g., client B) may include components that are identical or similar to the components of the electronic device (201) described in FIG. 2.
[0102] According to one embodiment, the first electronic device (201) may receive voice information in a first language according to the first user's speech during a video call or conversation, and transmit the received voice information in the first language to the server (1103). The first electronic device (201) may output the voice information in the second language received from the second electronic device (1101) or the server (1103) during the video call or conversation through the speaker (270), and may receive a translation (e.g., a first text in the first language) of the converted text (e.g., a second text) of the voice information in the second language and / or an additional description of a selected entity (e.g., an entity whose meaning is unknown in the first language) from the second electronic device (1101) or the server (1103), and may provide the received translation and / or the additional description (e.g., output through the speaker (270) or display through the display (260).
[0103] According to one embodiment, the second electronic device (1101) may receive voice information in a second language according to the second user's speech during a video call or conversation, and transmit the received voice information in the second language to the server (1103). The second electronic device (1101) may output the voice information in the first language received from the first electronic device (201) or the server (1103) during the video call or conversation through the speaker (270), and may receive a translation (e.g., a second text) of a text (e.g., a first text) converted from the voice information in the first language and / or an additional description of a selected entity (e.g., an entity whose meaning is unknown in the second language) from the first electronic device (201) or the server (1103), and may provide (e.g., output through the speaker or display through the display) the received translation and / or the additional description.
[0104] According to one embodiment, when the server (1103) receives speech information in a first language from the first electronic device (201) via the dialogue module (e.g., circuit or component) (301), the server (1103) may perform an operation to provide a translation (e.g., second text) of the first text of the speech information in the first language and / or an additional description. A specific description of each component of the dialogue module (301) executed in the server (1103) is substantially the same as the description in FIG. 3 and therefore is omitted.
[0105] According to one embodiment, the server (1103) may receive voice information in a first language from the first electronic device (201) via the conversation module (301) and convert the received voice information in the first language into text in the first language (e.g., first text).
[0106] According to one embodiment, the server (1103) can verify the context of the first text in the first language based on the conversation history accumulated in the history database (361) through the context verification module (320) by the conversation module (301), and translate the first text in the first language into the second text in the second language.
[0107] According to one embodiment, the server (1103) can identify at least one entity in a first language corresponding to at least one word in a first text by the dialogue module (301), and can identify at least one entity in a second language corresponding to at least one word in a second text. According to one embodiment, the server (1103) can compare at least one entity in the first language and at least one entity in the second language by the dialogue module (301), select (e.g., extract, confirm, or identify) an entity (e.g., a first entity) requiring additional explanation from among the entities in the at least one first language, obtain (or generate) data stored in a first entity database for the selected entity (e.g., the first entity) (or the original first entity) (e.g., “title,” “domain,” and “description” information), translate the obtained first description information into a second language, and refine the translated description information to generate (e.g., obtain) description information in the second language. Here, the first entity name may be a first language entity name (e.g., object or entity) for a word (e.g., slang and / or neologism) included in the first text whose meaning is unknown to the second language.
[0108] According to one embodiment, the server (1103) may transmit result information including a second text in a second language translated from a first text by the conversation module (301) and the acquired second description information in the second language to the second electronic device (1101). For example, the server (1103) may generate a user interface (UI) reflecting the result information including the second text and the acquired second description information in the second language, and transmit the generated UI to the second electronic device (1101).
[0109] According to one embodiment, the second electronic device (1101) may display a UI reflecting result information including second text in a second language and second description information in the acquired second language on a display (e.g., an execution screen displayed on the display).
[0110] According to one embodiment, when receiving voice information in a second language according to a second user's speech during a video call or conversation from a second electronic device (1101), the server (1103) may perform an operation to provide a translation (e.g., text in the first language) of the voice information in the second language received by the conversation module (301) and / or an additional description in the first language of a selected entity name (e.g., entity name whose meaning is unknown in the first language).
[0111] According to one embodiment, the server (1103) may convert voice information in a second language received from a second electronic device (1101) by the conversation module (301) into text (e.g., second text) in the second language, check the context for the entire conversation including the converted second language, and translate the second text in the second language into first text in the first language.
[0112] According to one embodiment, the server (1103) can identify at least one second language entity corresponding to at least one word in a second text by the dialogue module (301), and can identify at least one first language entity corresponding to at least one word in a translated first text. According to one embodiment, the server (1103) can compare at least one first language entity and at least one second language entity by the dialogue module (301), select (e.g., extract, confirm, or identify) an entity requiring additional explanation from the at least one second language entity, obtain (or generate) data stored in a second entity database for the selected entity (or original first entity) (e.g., “title,” “domain,” and “description” information) as second description information, translate the obtained second description information into a first language, and refine the translated description information to generate (e.g., obtain) first description information in the first language. Here, the selected entity name may be a second language entity name (e.g. entity or named entity (NE)) for a word (e.g. slang and / or neologism) contained in a second text of the second language whose meaning is unknown by the first language.
[0113] According to one embodiment, the server (1103) may transmit result information including a first text in a first language translated from a second text in a second language by the conversation module (301) and the acquired first description information in the first language to the second electronic device (1101). For example, the server (1103) may generate a UI reflecting the result information including the first text and the acquired first description information in the first language, and transmit the generated UI to the first electronic device (201).
[0114] According to one embodiment, the first electronic device (201) may display a UI reflecting result information including a first text in a first language and first description information in the first language obtained on a display (e.g., an execution screen displayed on the display).
[0115] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2) may include a communication circuit (e.g., communication module (190) of FIG. 1, communication circuit (230) of FIG. 2), a memory for storing instructions (e.g., memory (130) of FIG. 1, memory (220) of FIG. 2), and at least one processor (e.g., processor (120) of FIG. 1, processor (210) of FIG. 2).
[0116] According to one embodiment, the instructions, when executed by the at least one processor, may be configured to cause the electronic device to obtain a first text converted from speech information in a first language.
[0117] According to one embodiment, the instructions, when executed by the at least one processor, may be configured to cause the electronic device to generate a second text that is a translation of the first text into a second language.
[0118] According to one embodiment, the instructions, when executed by the at least one processor, may be configured to cause the electronic device to identify a first entity in the first text that requires additional description in the second language.
[0119] According to one embodiment, the instructions, when executed by the at least one processor, may be configured to cause the electronic device to obtain first description information that describes the first entity name in the first language.
[0120] According to one embodiment, the instructions, when executing the at least one processor, may be configured to cause the electronic device to generate second description information in the second language for the additional description based on the first description information.
[0121] According to one embodiment, the instructions may be configured to control the communication circuit to cause the electronic device to transmit the second text and the second description information to an external electronic device (e.g., the electronic device (104) of FIG. 1, the second electronic device (1101) of FIG. 11) when executing the at least one processor.
[0122] According to one embodiment, the instructions may be configured to cause the electronic device, when executing the at least one processor, to select a first model related to the first text from among the classified context models based on a previous conversation history, add a first context of the selected first model to the first text, and translate the first text into the second language based on the first context.
[0123] According to one embodiment, the instructions may be configured to cause the electronic device, when executing the at least one processor, to obtain at least one first language entity corresponding to at least one word included in the first text from a first entity database, to obtain at least one second language entity corresponding to at least one word included in the second text from a second entity database, to compare the at least one first language entity and the at least one second language entity, and to identify an entity among the at least one first language entity that does not match the at least one second language entity based on a result of the comparison.
[0124] According to one embodiment, the instructions may be configured to cause the electronic device, when executing the at least one processor, to obtain a similarity value of the at least one first language name based on a vector value of the at least one first language name and a vector value of the at least one second language name using a designated deep learning model, identify a first language name having a similarity value greater than or equal to a threshold value among the at least one first language name as a similar name, and identify that no additional description is required for the identified similar name, identify a first language name having a similarity value less than the threshold value among the at least one first language name as a dissimilar name, and identify that an additional description is required for the identified dissimilar name.
[0125] According to one embodiment, the instructions, when executing the at least one processor, may be configured to cause the electronic device to identify, as the first entity name, an entity name in the at least one first language that requires the additional description and is not searched for in the second entity name database.
[0126] According to one embodiment, the instructions may be configured to control the communication circuitry to, when executing the at least one processor, cause the electronic device to transmit the second text in the second language to the external electronic device based on the first entity name not being identified in the first text.
[0127] According to one embodiment, the system may further include a display (e.g., display (160) of FIG. 1, display (260) of FIG. 2) operatively connected to the at least one processor.
[0128] According to one embodiment, the instructions may be configured to cause the electronic device, when executing the at least one processor, to control the display to display an execution screen for a video call (e.g., a video call or a conversation), to control the display to display a screen of a first user speaking in the first language and a screen of at least one conversation partner user on the execution screen, and to control the display to display result information including the second text and / or the second description information on the screen of the at least one conversation partner user.
[0129] According to one embodiment, the instructions may be configured to cause the electronic device, when executing the at least one processor, to receive, from the external electronic device, result information including a fourth text in the first language that is a translation of the third text in the second language and description information in the first language for an entity name requiring additional description in the first language in the third text, and to control the display to display the received result information on a screen of the first user.
[0130] According to one embodiment, the system may further include a microphone operatively connected to at least one processor (e.g., input module (150) of FIG. 1, microphone (250) of FIG. 2).
[0131] According to one embodiment, the instructions, when executing the at least one processor, may be configured to cause the electronic device to receive the speech of the first user as the voice information in the first language through the microphone, and to convert the voice information in the first language into the first text.
[0132] According to one embodiment, the instructions are configured to cause the electronic device, when executing the at least one processor, to transmit a first text to a server (108, 1103) to confirm the first entity name, and to receive data of the first entity acquired from the first entity database from the server based on the server identifying, as the first entity name, among the at least one first language entity name that requires the additional description and is not searched for in the second entity database, and to obtain the first description information of the first language based on the data of the first entity, wherein the first entity database can store learned entity names of the first language based on a designated deep learning method, and the second entity database can store learned entity names of the second language based on the designated deep learning method.
[0133] FIG. 12 is a drawing showing an example of an operating method in an electronic device according to one embodiment, FIG. 13 is a drawing showing an example for a multi-party call in an electronic device according to one embodiment, and FIG. 14 is a drawing showing an example for a multi-party call in an electronic device according to one embodiment.
[0134] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0135] Referring to FIGS. 12 to 14, in operation 1201, an electronic device (201) according to an embodiment (e.g., the electronic device (101) of FIG. 1 and the electronic device (201) of FIGS. 2 and 11) can obtain voice information (e.g., a voice signal) (1303) of a first language spoken by a user (e.g., user A) during a multi-talker call (e.g., a video call or conversation) (e.g., “I think IU and Russemble’s songs are good these days.”). For example, as illustrated in FIG. 13, an electronic device (201) is an electronic device of a first user (user A) set to a first language (e.g., Korean), and the electronic device (201) performs a communication connection for a video call or conversation with an electronic device of a second user (user B) set to a second language (e.g., English) and an electronic device of a third user (user C) set to a third language (e.g., Japanese), and the electronic device (201) performs a video call or conversation among the first user, the second user, and the third user. According to one embodiment, when a video call or conversation is requested, as illustrated in FIG. 13, the electronic device (201) may execute a video call or conversation by a conversation module (e.g., the conversation module (301) of FIG. 3) and display an execution screen (1301) for the video call or conversation on the display (260). The electronic device (201) can display user screens (1310, 1320, 1330) for each user on the execution screen (1301). As illustrated in FIG. 13, the electronic device (201) can obtain voice information (1303) (e.g., “I think IU and Russemble’s songs are good these days”) of a first language (e.g., Korean) spoken by a first user (e.g., user A).
[0136] In operation 1203, the electronic device (201) according to one embodiment may convert voice information (1303) of a first language into a first text of the first language, and generate a second text translated from the converted first text into a second language. In one embodiment, the electronic device (201) may generate a second text translated from the first text of the first language (e.g., “I think IU and Loossemble songs are good these days”) into a second language (e.g., English) of a second user, as illustrated in FIG. 14. For example, as a third user participates in a video call or conversation, the electronic device may further generate a third text translated into a third language (e.g., “Recently, IU and Loossemble songs are good to think.”) in operation 1203.
[0137] In operation 1205, an electronic device (201) according to one embodiment may identify a first named entity (e.g., an entity or object) in a first text that requires additional explanation in a second language (e.g., the meaning of which is unknown in the second language). According to one embodiment, the electronic device (201) may identify (e.g., identify or extract) at least one named entity (e.g., “IU” and “Russemble”) in the first language corresponding to at least one word included in the first text, and may identify (e.g., select or identify) a first named entity (e.g., “Russemble”) requiring additional explanation among the identified at least one named entity in the first language. According to one embodiment, the electronic device (201) may use a first method (e.g., a rule-based method) to identify (e.g., identify or extract) at least one first language named entity from a first entity database (e.g., the NE DB (367) of FIGS. 3 and 11 ), identify (e.g., identify or extract) at least one second language named entity for at least one word included in a second text from a second entity database (e.g., the NE DB (367) of FIGS. 3 and 11 ), and compare the identified at least one first language named entity (e.g., a translated named entity) with the identified at least one second language named entity. The electronic device (201) may select a first language named entity (e.g., “Russemble”) that does not match at least one second language named entity among the at least one first language named entities. Here, the selected first language entity name may be an entity name that is included in the first entity name database and the translated entity name may not be searched in the second entity name database.According to one embodiment, the electronic device (201) may use a second method (e.g., a deep-based or deep learning based algorithm) to check the similarity between at least one first language entity and at least one second language entity (or entity information) and select the first language entity whose confirmed similarity value is higher than a pre-specified threshold as an entity that does not require additional explanation (e.g., “IU”). The electronic device (201) may select the first language entity whose confirmed similarity value is lower than the pre-specified threshold as an entity that requires additional explanation (e.g., “Russemble”). According to one embodiment, the electronic device (201) may check the first language entity (e.g., “Russemble”) selected using the first method and selected using the second method as the first entity. According to one embodiment, the electronic device (201) can use the first method and the second method to determine that at least one entity name in a first language (e.g., “IU” and “Russemble”) can be understood in a third language (e.g., Japanese) and that no additional information is required.
[0138] In operation 1207, an electronic device (201) according to an embodiment may obtain first description information in a first language for additional description of a first entity name in the identified first language. The electronic device may obtain (e.g., determine or generate) data of the identified first entity name in a first entity name database (e.g., "title": "Russemble", "domain": "singer", and "description": "a five-member girl group active in Korea") as the first description information.
[0139] In operation 1209, the electronic device (201) according to one embodiment may generate (e.g., obtain) second description information in a second language based on the first description information. The electronic device (201) may translate the first description information into the second language and refine the translated description information to generate second description information in the second language (e.g., "*Loossemble: It is a five-member girl group active in Korea."). According to one embodiment, as illustrated in FIG. 14, the electronic device (201) may not generate description information in a third language when, for example, the entity name "Loossemble" is detected in a third language (e.g., Japanese) database of a third language entity whose meaning can be understood.
[0140] In the 1211 operation, an electronic device (201) according to an embodiment may transmit result information (1401) including a second text in a second language (e.g., a translated text of the first text) and second description information in the second language to an external electronic device (e.g., the first external electronic device or the electronic device of a second user (User B)). For example, as shown in FIG. 14, the electronic device (201) may display the result information (1401) on a second user screen (1320) of an execution screen (1301). Here, the result information may include a second text (e.g., "I think IU and Loossemble songs are good these days.") obtained by translating a first text in a first language (e.g., "These days, I think IU and Loossemble songs are nice.") into the second language of the second user (e.g., English) and second description information in the second language (e.g., "*Loossemble: It is a five-member girl group active in Korea."). For example, as shown in FIG. 14, in the 1211 operation, as the third user participates in a video call or conversation and it is confirmed that there is no entity name for which additional explanation in the third language is required, a third text (e.g., 最近、IUとルセンブルの歌がいいと思います。) translated into the third language (e.g., Japanese)) can be transmitted to an electronic device (e.g., a second external electronic device) of a third user (user C). For example, the electronic device (201) can display the result information (1403) on the third user screen (1330) of the execution screen (1301), as illustrated in FIG. 14. The first user screen (1310), the second user screen (1320), and the third user screen (1330) displayed on the execution screen (1301) of the electronic device (201) can be synchronized with the user screens displayed on the execution screens of the second user's electronic device and the third user's electronic device, and can be displayed identically to the first user screen (1310), the second user screen (1320), and the third user screen (1330) displayed on the execution screen (1301). However, the electronic devices of each user can display only user screens for displaying only their own user.
[0141] According to one embodiment, if the electronic device (201) does not identify a first entity in the first text that requires additional explanation in a second language (e.g., the meaning of which is unknown in the second language), in operation 1211, the electronic device (201) may transmit result information that includes only a second text that is a translation of the first text into the second language to an external electronic device (e.g., an electronic device of a second user and / or an electronic device of a third user).
[0142] According to one embodiment, the electronic device (201) can transmit voice information (1303) of a first language received during a video call or conversation to an external electronic device, and when transmitting the voice information (1303) of the first language (e.g., in synchronization with the voice information of the first language), the electronic device can transmit the result information (1401 or 1403) to the external electronic device together.
[0143] Although the operation method of FIG. 12 as described above has been described as being performed in the electronic device (201), it is not limited thereto, and the operation method of FIG. 12 may be substantially identically performed by a server (e.g., server (1103) of FIG. 11). For example, when an external electronic device (e.g., second electronic device (1101) of FIG. 11) includes a conversation module (301) of FIG. 3, the external electronic device (e.g., electronic device of a second user or electronic device of a third user) may receive voice information (1303) of a first language spoken by the electronic device (201), and perform operations 1201 to 1209 of FIG. 12 in a substantially identical manner, and then display result information (1401 or 1403) including second text and second description information on an execution screen for a video call or conversation displayed on the display of the external electronic device.
[0144] According to one embodiment, the electronic device (201) may display, for example, the screen (1310) of the first user participating in a video call or conversation, the screen (1320) of the second user, and the screen (1330) of the third user on the execution screen (1301), as illustrated in FIG. 14. Without being limited thereto, the electronic device (201) may display the screen (1320) of the second user and / or the screen (1330) of the third user, who is the conversation partner, on the execution screen (1301). Without being limited to the screen examples illustrated in FIGS. 13 and 14, for example, the electronic device (201) may display only the screen (1310) of the first user on the execution screen (1301), or may display the screen (1320) of the second user and / or the screen (1330) of the third user without displaying the screen (1310) of the first user on the execution screen (1301). For example, the electronic device (201) can display at least one of the first user's screen (1310), the first user's screen (1310), or the third user's screen (1330) on the execution screen (1301) according to a user request or conditions set by the user using a designated menu (e.g., a button, an icon, or a graphic object).
[0145] The electronic device (201) may, for example, display result information (1401) including a translation and / or additional description in the corresponding language on the screen (1320) of the second user, as illustrated in FIG. 14, and / or display result information (1403) including a translation and / or additional description in the corresponding language on the screen (1330) of the third user. Without being limited thereto, the electronic device (201) may not display the translation and / or additional description in the corresponding language on the screen (1320) of the second user, who is the conversation partner, and / or the screen (1330) of the third user, respectively. For example, the translation and / or additional description in the second language may be displayed on an execution screen displayed on a display of the electronic device of the second user (e.g., the first external electronic device), and the translation and / or additional description in the third language may be displayed on an execution screen displayed on a display of the electronic device of the third user (e.g., the second external electronic device).
[0146] According to one embodiment, when an external electronic device (e.g., the second electronic device (1101) of FIG. 11) receives voice information in a second language according to the speech of a second user during a multi-talker call, the electronic device receives the voice information in the second language from the external electronic device or a server, receives result information including a translation (e.g., a first text in the first language) of the converted text (e.g., a second text) of the voice information in the second language and / or additional description (e.g., description information in the first language) of a selected entity (e.g., an entity whose meaning is unknown by the first language), and when outputting the voice information in the second language (e.g., outputting through a speaker), the electronic device may provide (e.g., output through a speaker or display through a display) result information including the received translation and / or additional description. For example, the electronic device may display a UI generated based on the translation and / or additional description or display the generated UI by reflecting it on an execution screen.
[0147] FIG. 15 is a diagram illustrating an example of a multi-party call in an electronic device according to one embodiment.
[0148] Referring to FIG. 15, an electronic device (201) according to one embodiment (e.g., the electronic device (101) of FIG. 1 and the electronic device (201) of FIG. 2) is an electronic device of a first user (user A) set to a first language (e.g., Korean), an electronic device of a second user (user B) set to a second language (e.g., English), and an electronic device of a third user (user C) set to a third language (e.g., Japanese) for performing a communication connection for a multi-talker call, and the electronic device (201) can be described as performing a multi-talker call among the first user, the second user, and the third user.
[0149] According to one embodiment, the electronic device (201) may perform a multi-party call (e.g., a video call or conversation) between a first user, a second user, and a third user by a conversation module (e.g., the conversation module (301) of FIG. 3), as illustrated in FIG. 15, and may display an execution screen (1301) for the video call or conversation, including user screens (1310, 1320, 1330) of each of the first user, the second user, and the third user, on the display (260). Without being limited thereto, the electronic device (201) may display only the first user's screen (1310) on the execution screen (1301).
[0150] According to one embodiment, the electronic device (201) may receive voice information (1501) of a third language (e.g., Japanese) spoken by a third user (e.g., user C) during a multi-talk call from a server or a third electronic device, as illustrated in FIG. 15.
[0151] According to one embodiment, when receiving voice information (1501) in a third language, as illustrated in FIG. 15, the electronic device (201) may receive result information (1503) including a translation (e.g., text in the first language) and / or description information for additional explanation (e.g., description information in the first language) for the voice information (1501) in the third language from a server, and display the result information (1503) including the received translation (e.g., text in the first language) and / or description information for additional explanation (e.g., description information in the first language) on the screen (1310) of the first user. The electronic device (201) may output the received voice information (1501) in the third language through a speaker. According to one embodiment, the server may convert voice information (1501) in a third language into third text (e.g., text in the third language) based on the operations of the video call module (301) described in FIG. 3, and translate the converted third text into other languages (e.g., a first language (Korean) and / or a second language (English)) set in electronic devices connected to the multi-talker call. The server may transmit result information (1503) including a translation (e.g., text in the first language) translated into the first language (e.g., Korean) and / or an additional description (e.g., description information in the first language) to the electronic device (201) (e.g., the first electronic device). The server may transmit result information (1505) including a translation (e.g., text in the second language) translated into the second language and an additional description (e.g., description information in the second language) to the second electronic device. The server can identify the entity name "Shinjuku" and the entity name "Yueni" as entity names that need to provide additional descriptions in third text of a third language.For example, the entity name for “Shinjuku” is an entity name whose meaning cannot be known by a second language and may not be detected in the second entity database, and the entity name for “Yueni” is an entity name whose meaning cannot be known by both the first language and the second language and may not be detected in the first entity database and the second entity database. The server may generate description information in the second language as an additional description for the entity name for “Shinjuku” and generate description information in the first language and description information in the second language as additional descriptions for the entity name for “Yueni.” In this way, the operation of providing a translation and additional description for the speech information (1501) in the third language may be performed substantially in the same manner as the operations of the server described above by receiving the speech information (1501) in the third language from the electronic device (201) including the conversation module (301) and / or the second electronic device, respectively. For example, if the electronic device (201) is a receiving device that receives voice information in a third language, a translation and additional description (e.g., description information in the first language) in a first language for a set language may be generated, and result information (1503) including the generated translation and additional description (e.g., description information in the first language) in the first language may be displayed on a first user screen (1310) displayed on a display (260) of the electronic device (201). For example, if the second electronic device is a receiving device that receives voice information (1501) in a third language, a translation and additional description (e.g., description information in the second language) in a second language for a set language may be generated, and the generated translation and additional description (e.g., description information in the second language) in the second language may be displayed on a second user screen displayed on a display of the second electronic device.For example, the server may generate result information including both translations (e.g., text in the first language and text in the second language) and / or additional descriptions (e.g., descriptive information in the first language and descriptive information in the second language) for each of the other languages (e.g., first language (Korean) and / or second language (English)) set on the electronic devices connected to the multi-language call, and provide the result information to the electronic device (201) (e.g., first electronic device) and the second electronic device, respectively. Accordingly, the electronic device (201) may display the UI provided on the execution screen (1301) (e.g., UI corresponding to the execution screen (1301) or UI corresponding to the first to third user screens (1310, 1320, 1330)), as illustrated in FIG. 15 . Without being limited thereto, the electronic device (201) may display only result information (1503) (e.g., UI) including a translation (e.g., text in the first language) and / or additional description (e.g., description information in the first language) on the first user's screen (1310) without displaying result information (1505) (e.g., UI) including a translation and / or additional information on the screens (1320 and 1330) of other users. For example, as illustrated in FIG. 15, the electronic device (201) can display result information (1503) including a translation of a third text of speech information (e.g., a speech signal) (1501) in a third language into a first language (e.g., “I thought of a girl group, and I saw Yueni at Shinjuku Station yesterday on my way home from work. We even took a picture together.”) and description information in the first language (e.g., “Yueni: A three-member under-idol active in Japan”) on the screen (1310) of the first user. The electronic device (201) can display a translation in a second language (English) (e.g., “A girl group came to mind, and I saw Yueni at Shinjuku Station yesterday on my way home from work. We even took a picture together.”).”) and the result information (1505) including the description information in the second language (e.g., “yueni: they are a three-member under-idol group active in Japan” and “Shinjuku: it is one of the largest cities located in Tokyo Japan.”) can be displayed on the second user’s screen (1320). Not limited to the screen example illustrated in FIG. 15, the electronic device (201) may display only the first user’s screen (1310) on the execution screen (1301) for a video call or conversation, or may display at least one of the first user’s screen (1310), the first user’s screen (1310), or the third user’s screen (1330) on the execution screen (1301) for a video call or conversation according to a user request or a condition set by the user by using a designated menu (e.g., a button, an icon, or a graphic object).
[0152] FIG. 16 is a diagram illustrating an example of a multi-party call in an electronic device according to one embodiment.
[0153] Referring to FIG. 16, an electronic device (201) according to an embodiment (e.g., the electronic device (101) of FIG. 1 and the electronic device (201) of FIG. 2) may obtain result information (1603) including a first translation (e.g., text in the first language) for voice information (e.g., “this is my first time hearting about under idol”) (1601) in a second language spoken by a second user, and display the result information (1603) on a first user screen (1310). Here, the example illustrated in FIG. 16 may be an example of obtaining result information (1603) without attempting to detect named entities (NEs) of language A, language B, and language C in the voice information (1601) in the second language spoken by the second user. The electronic device (201) or the server determines that an entity requiring additional explanation has not been identified, and may display a translation (e.g., text in the first language) of the second language voice information (1601) (e.g., "This is the first time I've ever heard of an under-idol") and a translation (e.g., text in the third language) (e.g., "Anda-aidol wa shimenezai senmaremashita.") in the third language on the first user's screen (1310) and the third user's screen (1330), respectively. Without being limited thereto, the electronic device (201) may display only a translation (e.g., text in the first language) of the second language voice information (1601) in the first language (e.g., "This is the first time I've ever heard of an under-idol") on the first user's screen (1310). Not limited to the screen example illustrated in FIG. 16, the electronic device (201) may display only the first user's screen (1310) on the execution screen (1301) for a video call or conversation, or may display at least one of the first user's screen (1310), the first user's screen (1310), or the third user's screen (1330) on the execution screen (1301) for a video call or conversation according to a user request or conditions set by the user using a designated menu (e.g., a button, an icon, or a graphic object).
[0154] Figure 17 is a diagram illustrating an example of an operating method in an electronic device according to one embodiment. In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0155] Referring to FIG. 17, in operation 1701, an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1 and the electronic device (201) of FIGS. 2 and 11) may obtain voice information (e.g., a voice signal) of a first language (e.g., Korean) spoken by a user (e.g., a speaker).
[0156] In operation 1703, an electronic device (201) according to one embodiment may convert speech information in a first language into first text in the first language.
[0157] In operation 1705, an electronic device (201) according to an embodiment may transmit a first text to a server (e.g., a server (1103) of FIG. 11) or an external electronic device (e.g., a second electronic device (1101) of FIG. 11). The server or the external electronic device may generate a text (e.g., a second text in the second language and / or a third text in the third language) that translates the first text in a first language into another language (e.g., a second language and / or a third language). The server or the external electronic device may detect (e.g., identify) a first entity (e.g., an entity or entity) in the first text that requires additional description in the second language (e.g., the meaning of which is unknown by the second language), and transmit data (e.g., “title”, “domain”, and “description” information) about the first entity to the electronic device as first description information. According to one embodiment, a server or an external electronic device may verify (e.g., identify or extract) at least one entity in a first language corresponding to at least one word included in a first text, and may verify (e.g., select or identify) a first entity requiring additional explanation among the verified at least one entity in the first language. According to one embodiment, the server or the external electronic device may verify (e.g., identify or extract) at least one entity in the first language from a first entity database using a first method (e.g., rule-based method), and may verify (e.g., identify or extract) at least one entity in a second language for at least one word included in a second text from a second entity database, and may compare the verified at least one entity in the first language (e.g., a translated entity) with the verified entity in the at least one second language. The server or the external electronic device may select a entity in the first language that does not match at least one entity in the second language among the entities in the at least one first language.Here, the selected first language entity name may be an entity name that is included in the first entity database and the translated entity name may be an entity name that is not searched for in the second entity database. According to one embodiment, the server or the external electronic device may use a second method (e.g., a deep-based or deep learning based algorithm) to check the similarity between at least one first language entity name (or entity information) and at least one second language entity name (or entity information), and select the first language entity name whose confirmed similarity value is higher than a pre-specified threshold as an entity name that does not require additional explanation. The server or the external electronic device may select the first language entity name whose confirmed similarity value is lower than a pre-specified threshold and is not similar (e.g., has low similarity), as an entity name that requires additional explanation. According to one embodiment, the server or the external electronic device may check the first language entity name that is selected using the first method and selected using the second method as the first entity.
[0158] In operation 1707, an electronic device according to one embodiment may receive first description information (e.g., data about the first entity) in a first language for additional description of the first entity in the identified first language from a server or an external electronic device.
[0159] In operation 1709, an electronic device according to one embodiment may generate (e.g., obtain) second description information in a second language based on first description information. The electronic device may translate the first description information into the second language and refine the translated description information to generate second description information in the second language.
[0160] In operation 1711, an electronic device according to an embodiment may transmit result information including a second text in a second language (e.g., a translation of the first text) and second descriptive information in the second language to an external electronic device. For example, operations 1707 to 1711 may be performed by a server or an external electronic device, and the electronic device may receive result information generated by the server or the external electronic device (e.g., result information including a second text in a second language (e.g., a translation of the first text) and second descriptive information in the second language) and display the result information on a display.
[0161] According to one embodiment, if the first text does not have a first entity name that requires additional description in the second language, the electronic device may not receive the first description information from the server, and in operation 1711, the electronic device may transmit result information including only the second text translated from the first text into the second language to the server or an external electronic device.
[0162] According to one embodiment, the electronic device can transmit voice information of a first language to an external electronic device during a multi-talk call, and when transmitting the voice information of the first language (e.g., in synchronization with the voice information of the first language), the electronic device can transmit the result information together with the voice information of the first language to the external electronic device.
[0163] FIG. 18 is a diagram illustrating an example of an operation method between electronic devices and a server according to one embodiment.
[0164] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0165] Referring to FIG. 18, in operation 1801, an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1 and the electronic device (201) of FIGS. 2 and 11) may acquire voice information (e.g., a voice signal) of a first language spoken by a first user (e.g., a speaker) (e.g., a third user), and in operation 1803, may transmit the acquired voice information of the first language to a server (1103).
[0166] In operation 1805, a server (1103) according to one embodiment may convert speech information in a first language into a first text in the first language and generate a second text by translating the converted first text into a second language. When generating the second text, the server (1103) may acquire a context and add the acquired context to the beginning of the first text to perform a more accurate translation using the context.
[0167] In operation 1807, a server (1103) according to one embodiment may detect (e.g., identify) a first entity (e.g., an object or entity) in a first text that requires additional explanation in a second language (e.g., whose meaning is unknown by the second language).
[0168] In operation 1809, a server (1103) according to one embodiment may obtain data for a first entity name (e.g., “title”, “domain”, and “description” information) from a first entity name database, and obtain first description information based on the obtained data for the first entity name.
[0169] In operation 1811, the server (1103) according to one embodiment may generate second description information in a second language for the acquired first description information. The server (1103) may translate the first description information into the second language and refine the translated description information to generate the second description information in the second language.
[0170] In operation 1813, the server (1103) according to one embodiment may transmit result information including second text in a second language (e.g., a translation of the first text) and second description information in the second language to the second electronic device (1101). For example, the server (1103) may also transmit the result information to the first electronic device (201) and display it on the second user's screen included in the execution screen displayed on the first electronic device (201) (e.g., the second user's screen (1320) of FIGS. 13 to 16).
[0171] In operation 1815, the second electronic device (1101) according to one embodiment may display result information including second text in a second language (e.g., a translation of the first text) and second description information in the second language on an execution screen displayed on a display of the second electronic device (1101). According to one embodiment, the second electronic device (1101) may receive voice information in a first language from the first electronic device (201) through the server (1103) during a multi-talker call, and when receiving and outputting the voice information in the first language (e.g., in synchronization with the voice information in the first language), the second electronic device (1101) may also receive and display the result information from the server (1103).
[0172] According to one embodiment, if the server (1103) does not have a first entity name that requires additional description in the second language in the first text, the electronic device (201) and the second electronic device (1101) may each receive result information including only a translation (e.g., a second text in the second language) without receiving the first description information from the server (1103).
[0173] FIG. 19 is a diagram illustrating an example of an operating method in a second electronic device according to one embodiment. In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0174] Referring to FIG. 19, in operation 1901, a second electronic device (e.g., electronic device (104) of FIG. 1 and electronic device (1101) of FIG. 11) according to one embodiment may receive voice information (e.g., voice signal) of a first language (e.g., Korean) spoken by a first user (e.g., speaker) from a first electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2 and FIG. 11).
[0175] In operation 1903, a second electronic device according to one embodiment may convert speech information in a first language into a first text in the first language, and generate text (e.g., a second text in the second language and / or a third text in the third language) translated from the converted first text into another language (e.g., a second language and / or a third language).
[0176] In operation 1905, a second electronic device according to an embodiment may identify (e.g., detect) a first entity (e.g., an entity or object) in a first text that requires additional explanation in a second language (e.g., whose meaning is unknown in the second language). The second electronic device according to an embodiment may identify (e.g., identify or extract) at least one entity in the first language corresponding to at least one word included in the first text, and identify (e.g., select or identify) a first entity requiring additional explanation among the identified at least one entity in the first language. A second electronic device according to one embodiment may use a first method (e.g., a rule-based method) to identify (e.g., identify or extract) at least one first language named entity from a first entity database, identify (e.g., identify or extract) at least one second language named entity for at least one word included in a second text from a second entity database, and compare the identified at least one first language named entity (e.g., a translated named entity) with the identified at least one second language named entity. The second electronic device according to one embodiment may select a first language named entity that does not match at least one second language named entity among the at least one first language named entities. Here, the selected first language named entity may be an entity that is included in the first entity database, and whose translated named entity is not searched for in the second entity database. A second electronic device according to one embodiment may use a second method (e.g., a deep-based or deep learning based algorithm) to determine the similarity between at least one first language entity name (or entity name information) and at least one second language entity name (or entity name information), and select an entity name of the first language whose determined similarity value is greater than or equal to a predetermined threshold value as an entity name that does not require additional explanation.According to one embodiment, a second electronic device may select a first language entity whose identified similarity value is less than a predetermined threshold (e.g., has low similarity) as a first language entity requiring additional explanation. According to one embodiment, a second electronic device may select a first language entity selected using the first method and identify the entity selected using the second method as a first language entity.
[0177] In operation 1907, a second electronic device according to one embodiment may obtain data for a first entity name (e.g., “title”, “domain”, and “description” information) as first description information.
[0178] In operation 1909, a second electronic device according to one embodiment may generate (e.g., obtain) second description information in a second language based on the first description information. The electronic device may translate the first description information into the second language and refine the translated description information to generate second description information in the second language.
[0179] In operation 1911, the second electronic device according to one embodiment may display result information including second text in a second language (e.g., a translation of the first text) and second descriptive information in the second language.
[0180] According to one embodiment, if the first text does not have a first entity name that requires additional description in the second language, the second electronic device may not obtain the first description information, and in operation 1911, the second electronic device may display result information that includes only second text that is a translation of the first text into the second language.
[0181] According to one embodiment, the second electronic device may transmit voice information of a first language to the first electronic device during a multi-talk call, and when outputting the voice information of the first language (e.g., in synchronization with the voice information of the first language), the second electronic device may display the result information.
[0182] According to one embodiment, a method of operating in an electronic device (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2 and FIG. 11) may include an operation of obtaining a first text converted from speech information of a first language.
[0183] According to one embodiment, a method of operation in an electronic device may include generating a second text translated from the first text into a second language.
[0184] According to one embodiment, a method of operation in an electronic device may include identifying a first entity in the first text that requires additional description in the second language.
[0185] According to one embodiment, a method of operating in an electronic device may include obtaining first description information that describes the first entity name in the first language.
[0186] According to one embodiment, the method of operation in the electronic device may include an operation of generating second description information in the second language for the additional description based on the first description information.
[0187] According to one embodiment, the method of operation in the electronic device may include an operation of transmitting the second text and the second description information to an external electronic device (e.g., the electronic device (104) of FIG. 1, the second electronic device (1101) of FIG. 11).
[0188] According to one embodiment, the operation of identifying a first entity requiring additional explanation in the second language may include an operation of selecting a first model related to the first text from among classified context models based on a previous conversation history, and an operation of adding a first context of the selected first model to the first text and translating the first text into the second language based on the first context.
[0189] According to one embodiment, the operation of identifying a first entity requiring additional explanation in the second language may include an operation of obtaining at least one entity in the first language corresponding to at least one word included in the first text from a first entity database, an operation of obtaining at least one entity in the second language corresponding to at least one word included in the second text from a second entity database, and an operation of comparing the entity in the at least one first language and the entity in the at least one second language to identify an entity among the entity in the at least one first language that does not match the entity in the at least one second language.
[0190] According to one embodiment, the operation of identifying a first entity requiring additional explanation in the second language may include an operation of comparing a similarity between a vector value of an entity in the at least one first language and a vector value of an entity in the at least one second language using a designated deep learning model, an operation of identifying an entity in the first language having a similarity value greater than or equal to a threshold value among the entities in the at least one first language as a similar entity and confirming that no additional explanation is required for the identified similar entity, and an operation of identifying an entity in the first language having a similarity value less than the threshold value among the entities in the at least one first language as a dissimilar entity and confirming that an additional explanation is required for the identified dissimilar entity.
[0191] According to one embodiment, the operation of identifying a first entity name requiring additional description in the second language may include an operation of identifying an entity name among the at least one entity name in the first language that requires additional description and is not searched in the second entity name database as the first entity name.
[0192] According to one embodiment, the method may further include transmitting the second text in the second language to the external electronic device based on the first entity requiring additional description in the second language not being identified in the first text.
[0193] According to one embodiment, the method may further include an operation of displaying an execution screen for a multi-party call on a display (160, 260) of the electronic device, an operation of displaying a screen of the first user and a screen of at least one conversation partner user on the execution screen, and an operation of displaying result information including the second text and / or the second description information on a screen of the at least one conversation partner user.
[0194] According to one embodiment, the method may further include an operation of receiving, from the external electronic device, result information including a fourth text in the first language that is a translation of the third text in the second language and description information in the first language for entity names in the third text that require additional description in the first language, and an operation of reflecting the received result information on the screen of the first user and displaying it on the display.
[0195] According to one embodiment, an electronic device (e.g., an electronic device (104) of FIG. 1, a second electronic device (1101) of FIG. 11) includes a display, a communication circuit, a memory storing instructions, and at least one processor, wherein the instructions, when executed by the at least one processor, cause the electronic device to receive voice information in a first language from the first electronic device, generate a second text by translating a first text in the first language into a second language by converting the received voice information, identify a first entity requiring additional description in the second language from the first text, obtain first description information that describes the first entity in the first language, generate second description information in the second language for the additional description based on the first description information, and control the display to display the second text in the second language and the second description information in the second language.
[0196] According to one embodiment, in a non-transitory storage medium storing a program, the program may include executable instructions that, when executed by an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIGS. 2 and 11) at least one processor (e.g., the processor (120) of FIG. 1, the processor (210) of FIG. 2), cause the electronic device to perform an operation of obtaining a first text converted from speech information in a first language, an operation of generating a second text translated from the first text into a second language, an operation of identifying a first entity requiring additional description in the second language from the first text, an operation of obtaining first description information describing the first entity in the first language, an operation of generating second description information in the second language for the additional description based on the first description information, and an operation of transmitting the second text and the second description information to an external electronic device (e.g., the electronic device (104) of FIG. 1, the second electronic device (1101) of FIG. 11).
[0197] An electronic device according to one embodiment of the present disclosure does not simply provide a translation into another language during a multi-talker call, but rather identifies the entity name of a new word or slang whose meaning cannot be understood by the other language contained in the text converted from the voice information, and provides additional explanations for the identified entity name, thereby enabling other users who speak other languages to accurately understand the meaning of the entity name and conduct a smooth conversation. In addition, various effects that can be directly or indirectly understood through the present disclosure may be provided. The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art to which the present disclosure pertains from the description below.
[0198] The embodiments disclosed in this disclosure are presented for the purpose of explaining and understanding the disclosed technical content, and do not limit the scope of the technology described in this disclosure. Accordingly, the scope of this disclosure should be interpreted to include all modifications or various other embodiments based on the technical concepts of this disclosure.
[0199] Electronic devices according to various embodiments disclosed in the present disclosure may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to embodiments of the present disclosure are not limited to the aforementioned devices.
[0200] The various embodiments of the present disclosure and the terminology used therein are not intended to limit the technical features described in the present disclosure to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In the present disclosure, each of the phrases "A or B," "at least one of A and B," "at least one of A or B," "A, B, or C," "at least one of A, B, and C," and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among the phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0201] The term "module" used in various embodiments of the present disclosure may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0202] Various embodiments of the present disclosure may be implemented as software (e.g., a program (140)) including one or more commands stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one command among the one or more commands stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one command called. The one or more commands may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0203] According to one embodiment, the method according to various embodiments disclosed in the present disclosure may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0204] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.< / ai> < / ai> < / context> < / context>
Claims
1. In an electronic device (101, 201), Communication circuit (190, 230); Memory (130, 220) for storing instructions; and comprising at least one processor (120, 210), The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtain a first text converted into speech information of the first language, Generate a second text by translating the first text into a second language, In the above first text, identify the first entity name that requires additional explanation in the second language, Obtain first description information that describes the first entity name in the first language, Based on the above first description information, second description information in the second language is generated for the additional description, An electronic device that controls the communication circuit to transmit the second text and the second description information to an external electronic device (104, 1101).
2. In the first paragraph, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the previous conversation history, select the first model related to the first text among the classified context models, An electronic device that adds a first context of the selected first model to the first text and translates the first text into the second language based on the first context.
3. In claim 1 or 2, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtaining at least one first language entity name corresponding to at least one word included in the first text from a first entity name database, Obtaining at least one second language entity name corresponding to at least one word included in the second text from a second entity name database, An electronic device that compares at least one first language entity and at least one second language entity, and, based on the comparison result, identifies an entity among the at least one first language entity that does not match the at least one second language entity.
4. In any one of paragraphs 1 to 3, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtaining a similarity value of at least one first language entity based on a vector value of at least one first language entity and a vector value of at least one second language entity using a designated deep learning model, Identifying a first language entity among the at least one first language entity having a similarity value greater than or equal to a threshold value as a similar entity, and identifying that no additional description is required for the identified similar entity, An electronic device that identifies a first language entity having a similarity value less than a threshold value among at least one first language entity as a dissimilar entity, and identifies the identified dissimilar entity as requiring additional description.
5. In any one of paragraphs 1 to 4, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Identifying an entity name among the entity names of at least one first language that requires the additional explanation and is not searched in the second entity name database as the first entity name, An electronic device that controls the communication circuit to transmit the second text in the second language to the external electronic device based on the first entity name not being identified in the first text.
6. In any one of paragraphs 1 to 5, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Controlling the display (160, 260) of the electronic device to display an execution screen for multi-talking, Controlling the display to display the screen of the first user speaking in the first language and the screen of at least one conversation partner user on the execution screen; An electronic device that controls the display to display result information including the second text and / or the second description information on a screen of the at least one conversation partner user.
7. In any one of paragraphs 1 to 6, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Receiving, from the external electronic device, result information including a fourth text in the first language that is a translation of the third text in the second language and explanatory information in the first language for entity names in the third text that require additional explanation in the first language; An electronic device that controls the display to display the received result information on the screen of the first user.
8. In any one of paragraphs 1 to 7, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: The speech of the first user is received as voice information of the first language through the microphone (150, 250) of the electronic device, Converting the speech information of the first language into the first text, and transmitting the first text to the server (108, 1103) to confirm the first entity name, By the server, based on identifying an entity name that requires additional description among the at least one first language entity name and is not searched for in the second entity name database as the first entity name, data of the first entity name obtained from the first entity name database is received from the server, Obtain the first description information of the first language based on the data of the first entity name, The above first entity name database stores learned entity names of the first language based on a designated deep learning method, An electronic device in which the second entity name database stores learned entity names of the second language based on the specified deep learning method.
9. In the method of operation in an electronic device (101, 201), An action of obtaining a first text converted into speech information of a first language; An action of generating a second text translated from the first text into a second language; An action to identify a first entity name in the first text that requires additional explanation in the second language; An operation of obtaining first description information that describes the first entity name in the first language; An operation of generating second description information in the second language for the additional description based on the first description information; and A method comprising the action of transmitting the second text and the second description information to an external electronic device (104, 1101).
10. In paragraph 9, the operation of confirming the first entity name requiring additional explanation in the second language is as follows: An operation of selecting a first model related to the first text from among the classified context models based on the previous conversation history; and A method comprising adding a first context of the selected first model to the first text and translating the first text into the second language based on the first context.
11. In the 9th or 10th paragraph, the operation of confirming the first entity name requiring additional explanation in the second language is as follows: An operation of obtaining at least one entity name in a first language corresponding to at least one word included in the first text from a first entity name database; An operation of obtaining at least one entity name in a second language corresponding to at least one word included in the second text from a second entity name database; and A method comprising the action of comparing at least one entity name of the first language and at least one entity name of the second language, and identifying an entity name among the at least one entity name of the first language that does not match an entity name of the at least one second language.
12. In any one of paragraphs 9 to 11, the operation of confirming the first entity name requiring additional explanation in the second language is as follows: An operation of comparing the similarity between a vector value of an entity name of at least one first language and a vector value of an entity name of at least one second language using a designated deep learning model; An operation of identifying a first language entity having a similarity value greater than or equal to a threshold value among at least one first language entity names as a similar entity name and confirming that no additional explanation is required for the identified similar entity name; and A method comprising the steps of identifying a first language entity having a similarity value less than a threshold value among at least one first language entity as a dissimilar entity, and determining that an additional description is required for the identified dissimilar entity.
13. In any one of paragraphs 9 to 12, the operation of confirming the first entity name requiring additional explanation in the second language is as follows: The method comprises an operation of identifying an entity name in at least one first language that requires additional explanation and is not searched for in the second entity name database as the first entity name, wherein: A method further comprising the action of transmitting said second text in said second language to said external electronic device based on the first entity name requiring additional description in said second language not being identified in said first text.
14. In any one of clauses 9 to 13, the method, An action of displaying an execution screen for a multi-party call on the display (160, 260) of the electronic device; An action of displaying the screen of the first user and the screen of at least one conversation partner user on the execution screen; An action of displaying result information including the second text and / or the second description information on a screen of the at least one conversation partner user; An operation of receiving, from the external electronic device, a fourth text in the first language that is a translation of the third text in the second language, and result information including first language description information for entity names in the third text that require additional description in the first language; and Further comprising an action of reflecting the received result information on the screen of the first user and displaying it on the display, The operation of obtaining the first text converted into the speech information of the first language is as follows: An operation of receiving the speech of the first user as voice information in the first language through the microphone; and A method comprising the action of converting said speech information of said first language into said first text.
15. In electronic devices (104, 1101), display; communication circuit; Memory for storing instructions; and comprising at least one processor, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Receive voice information of a first language from the first electronic device, Generate a second text translated from a first text in a first language into a second language by converting the received voice information, In the above first text, identify the first entity name that requires additional explanation in the second language, Obtain first description information that describes the first entity name in the first language, Based on the above first description information, second description information in the second language is generated for the additional description, An electronic device that controls the display to display the second text in the second language and the second descriptive information in the second language.
Citation Information
Patent Citations
Automatic voice translation system for setting translation language by voice input, automatic voice translation method, and program thereof
JP2021086404A
Translation method and electronic device
JP2023029846A
System for Speaker Diarization based Multilateral Automatic Speech Translation System and its operating Method, and Apparatus supporting the same
KR1020150093482A
Security Sticker
KR102055205B1
Method and system for remote communication based on real-time translation service
KR102264224B1