Electronic device for providing translation service and method performed thereby

The electronic device uses a neural network-based text modification model to enhance translation accuracy in noisy conditions by modifying input text data, addressing the challenges of unclear speech and environmental noise in translation technologies.

WO2026023826A1PCT designated stage Publication Date: 2026-01-29SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/007221
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-16
Filing Date
2025-05-28
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing translation technologies struggle with accuracy in noisy environments and unclear speech, leading to suboptimal communication outcomes.

Method used

An electronic device employs a neural network-based text modification model to modify input text data in noisy conditions, enhancing translation accuracy by generating modified text data that better fits the conversation context.

Benefits of technology

Improves translation accuracy in noisy environments by modifying input text data based on user-specific patterns and conversation context, ensuring clearer and more contextually appropriate translations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025007221_29012026_PF_FP_ABST
    Figure KR2025007221_29012026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device and a method performed by the electronic device are disclosed. The electronic device may comprise: a microphone for acquiring input speech data; a memory comprising one or more storage media for storing instructions; and one or more processors comprising a processing circuit. When the instructions are individually or collectively executed by the one or more processors, the instructions may instruct the electronic device to: convert input speech data in a first language acquired through the microphone into input text data in the first language corresponding to the input speech data; if the noise level of sound data inputted through the microphone satisfies a condition, generate, using a neural network-based text modification model, modified text data in the first language in which at least a portion of text in the input text data has been changed; and generate translation result text data in a second language on the basis of the modified text data.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for providing translation services and method performed thereby

[0001] The present disclosure relates to an electronic device providing a translation service and a method performed thereby.

[0002] Applications running on electronic devices (e.g., translation apps, voice assistants, messenger apps, or call apps) can provide translation (or interpretation) functionality. Sentences entered by a user through a user interface (UI) are translated using the translation function, and the translated results can be displayed on the electronic device's screen or output through the device's speakers.

[0003] The above information may be provided as background information to aid in understanding the present disclosure. None of the above is claimed to be prior art related to the present disclosure, nor can it be used to determine prior art.

[0004] An electronic device according to an embodiment may include a microphone for acquiring input voice data, a memory including one or more storage media for storing instructions, and one or more processors including a processing circuit. When the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to convert input voice data in a first language acquired through the microphone into input text data in a first language corresponding to the input voice data, generate modified text data in the first language in which at least a part of the text is changed in the input text data using a neural network-based text modification model when a noise level of sound data input through the microphone satisfies a condition, and generate translated result text data in a second language based on the modified text data.

[0005] A method performed by an electronic device according to an embodiment may include an operation of acquiring input voice data in a first language through a microphone of the electronic device. The method may include an operation of converting the input voice data into input text data in the first language corresponding to the input voice data. The method may include an operation of determining whether a noise level of sound data input through the microphone satisfies a condition. If it is determined that the noise level satisfies the condition, the method may include an operation of generating modified text data in the first language in which at least a portion of the text in the input text data is changed using a neural network-based text modification model. The method may include an operation of generating translated result text data in a second language based on the modified text data.

[0006] FIG. 1 is a block diagram illustrating an exemplary configuration of an electronic device according to various embodiments.

[0007] FIG. 2A is a block diagram illustrating an integrated intelligence system according to various embodiments.

[0008] FIG. 2b is a block diagram illustrating a configuration of an electronic device that provides a translation service according to various embodiments.

[0009] FIG. 3 is a diagram for explaining outputting voice output data according to text to speech (TTS) during translation during a call according to various embodiments.

[0010] FIG. 4 is a diagram for explaining an electronic device that outputs voice output data according to text-to-speech conversion during translation during a call according to various embodiments.

[0011] FIG. 5 is a diagram for explaining a translation service according to various embodiments.

[0012] FIG. 6 is a diagram illustrating an example of a method for determining the output of text-to-speech conversion during translation during a call according to various embodiments.

[0013] FIGS. 7 and 8 are flowcharts for explaining operations of a method for providing a translation service performed by an electronic device according to various embodiments.

[0014] FIG. 9 is a flowchart illustrating operations of a method performed by an electronic device providing a translation service according to various embodiments.

[0015] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.

[0016]

[0017] FIG. 1 is a block diagram illustrating an exemplary configuration of an electronic device according to various embodiments.

[0018] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with another electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of another electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108).

[0019] According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor (176), an interface (177), a connection terminal (178), a haptic module (179), a camera (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor (176), the camera (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0020] The processor (120) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing. The processor (120) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs (140), data, etc.) stored in the memory (130). The processor (120) may include a processor assembly including one or more processing circuits. The processor (120) may include any processing circuit that is operative to control the performance and operations of one or more components of the electronic device (101) (e.g., the memory (130), the display module (160), the camera (180), the communication module (190), and / or the sensor (176)).

[0021] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor (176) or a communication module (190)) in the volatile memory (132), process the commands or data stored in the volatile memory (132), and store the resulting data in the non-volatile memory (134).

[0022] According to one embodiment, the processor (120) may include one or more processors, and the operations of the electronic device (101) described in this disclosure may be performed by one processor or by a combination of multiple processors. As used herein, a “processor” may include a processing circuit or may include multiple processors. For example, as used herein, including in the claims, the term “processor” may include various processing circuits including one or more processors, wherein the one or more processors may be configured to perform various functions described in this disclosure, individually and / or collectively, in a distributed manner. When “processor,” “at least one processor,” and “one or more processors” are described herein as being configured to perform multiple functions, these terms include, but are not limited to, situations where one processor performs some of the recited functions and another processor performs other of the recited functions, and situations where a single processor may perform all of the recited functions. Furthermore, the one or more processors may include a combination of processors that perform various recited / disclosed functions, for example, in a distributed manner. One or more processors may execute instructions to accomplish or perform various functions.

[0023] According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). When the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0024] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0025] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0026] In one embodiment, the memory (130) may include one or more memories. Instructions for controlling the processor (120) to perform operations of the electronic device (101) described in the present disclosure may be stored in one memory or may be stored in multiple memories.

[0027] The program (140) may be stored as software in the memory (130). The program (140) may include, for example, an operating system (142), middleware (144), or an application (146).

[0028] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0029] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0030] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0031] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), or output sound through the sound output module (155), or another electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0032] The sensor (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor (176) may include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor. For example, the sensor (176) may include an inertial measurement unit (IMU).

[0033] The interface (177) may support one or more designated protocols that may be used to allow the electronic device (101) to connect directly or wirelessly with another electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0034] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to another electronic device (e.g., electronic device (102)). In one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0035] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0036] The camera (180) can capture still images and moving images. According to one embodiment, the camera (180) can include one or more lenses, one or more image sensors, one or more image signal processors, or one or more flashes.

[0037] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as at least a part of a power management integrated circuit (PMIC), for example.

[0038] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0039] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and another electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication circuits. The communication module (190) may include one or more communication processors (CPs) that operate independently from the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). These communication modules can communicate with external electronic devices (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data relation (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), another electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)).

[0041] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0042] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0043] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0044] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199).

[0045] Each of the external electronic devices, such as other electronic devices (102, 104) and the server (108), may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed by the electronic device (101) may be executed by one or more external electronic devices among the other electronic devices (102, 104) or the server (108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of executing the function or service itself or in addition, request one or more external electronic devices to execute at least a part of the function or service. The one or more external electronic devices that receive the request may execute at least a part of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a part of a response to the request.

[0046] FIG. 2a is a block diagram illustrating an integrated intelligent system according to various embodiments.

[0047] Referring to FIG. 2A, an integrated intelligent system (20) according to one embodiment may include an electronic device (201) (e.g., electronic device (101) of FIG. 1), an intelligent server (200) (e.g., server (108) of FIG. 1), and a service server (300) (e.g., server (108) of FIG. 1).

[0048] In one embodiment, the electronic device (201) may be, but is not limited to, a smartphone, a tablet personal computer, a mobile phone, a speaker (e.g., an AI speaker), a video phone, an e-book reader, a desktop personal computer, a laptop personal computer, a netbook computer, a workstation, a server, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, a wearable device, a virtual reality (VR) device, or an augmented reality (AR) device. The electronic device (201) may include a communication interface (202) (e.g., interface (177) of FIG. 1), a microphone (206) (e.g., input module (150) of FIG. 1), a speaker (205) (e.g., audio output module (155) of FIG. 1), a display module (204) (e.g., display module (160) of FIG. 1), a memory (207) (e.g., memory (130) of FIG. 1), or a processor (203) (e.g., processor (120) of FIG. 1). The above-listed components may be operatively or electrically connected to each other.

[0049] The communication interface (202) can be connected to an external device to transmit / receive data. The microphone (206) can receive sound (e.g., user speech) and convert it into an electrical signal. The speaker (205) can output the electrical signal as sound (e.g., voice). The display module (204) can display an image or video. The display module (204) can also display a graphical user interface (GUI) of an app (or application program) that is being executed. The display module (204) can receive a touch input through a touch sensor. For example, the display module (204) can receive a text input through a touch sensor in an on-screen keyboard area displayed within the display module (204).

[0050] The memory (207) can store a client module (209), a software development kit (SDK) (208), and multiple applications (e.g., a first app (211_1), a second app (211_2)). The client module (209) and the SDK (208) can configure a framework (or solution program) for performing general functions. In addition, the client module (209) or the SDK (208) can configure a framework for processing user input (e.g., voice input, text input, touch input).

[0051] In one embodiment, each of the multiple applications stored in memory (207) may be a program for performing a specified function. The multiple applications may include, for example, an alarm app, a call app, a translation app, a messaging app, and / or a scheduling app. According to one embodiment, the multiple applications may be executed by the processor (203) to sequentially execute at least some of the defined operations.

[0052] The processor (203) can control the overall operation of the electronic device (201). For example, the processor (203) can be electrically connected to a communication interface (202), a microphone (206), a speaker (205), and a display module (204) to perform a specified operation.

[0053] The processor (203) may execute a program stored in the memory (207) to control the electronic device (210) to perform a specified operation and / or function. For example, the processor (203) may execute at least one of the client module (209) or the SDK (208) to perform an operation for processing user input.

[0054] The client module (209) can receive user input. For example, the client module (209) can receive a voice signal (or input voice data) corresponding to a user utterance detected through the microphone (206). In addition, the client module (209) can receive a touch input detected through the display module (204) and / or a text input detected through a keyboard or a visual keyboard. In addition, the client module (209) can receive various types of user input detected through an input module connected to the electronic device (201). The client module (209) can transmit the received user input to the intelligent server (200). The client module (209) can transmit status information of the electronic device (201) to the intelligent server (200) together with the received user input. The status information can include, for example, execution status information of an application.

[0055] In one embodiment, the client module (209) can receive a result corresponding to the received user input. If the intelligent server (200) can produce a result corresponding to the received user input, the client module (209) can receive the result corresponding to the user input from the intelligent server (200). The client module (209) can display the received result on the display module (204). In addition, the client module (209) can output the received result as audio through the speaker (205).

[0056] In one embodiment, the client module (209) may receive a request from the intelligent server (200) to obtain information necessary to produce a result corresponding to a user input. The client module (209) may transmit the necessary information to the intelligent server (200) in response to the received request.

[0057] In one embodiment, the client module (209) may include a voice recognition module. The client module (209) may recognize a voice input to perform limited functions through the voice recognition module. The intelligent server (200) may receive information related to the user voice input from the electronic device (201) via a communication network. The intelligent server (200) may convert data related to the received voice input into text data. The intelligent server (200) may generate a plan for performing a task corresponding to the user voice input based on the text data. The plan may include information about parameters required for performing one or more operations and an execution order of one or more operations.

[0058] In one embodiment, the plan may be generated by an artificial intelligence (AI) system. The AI ​​system may be a rule-based system, a neural network-based system (e.g., a feedforward neural network (FNN), a recurrent neural network (RNN), or a combination of the foregoing or another AI system. The plan may be selected from a set of predefined plans or may be generated in real time in response to a user request.

[0059] The intelligent server (200) can transmit the processing result according to the generated plan to the electronic device (201) or transmit the generated plan to the electronic device (201). The electronic device (201) can display the processing result according to the plan on the display module (204). The electronic device (201) can display the result of executing an operation according to the plan on the display module (204).

[0060] In one embodiment, the intelligent server (200) may include a front end (215), a natural language platform (220), a capsule database (230), an execution engine (240), an end user interface (250), a management platform (260), a big data platform (270), and an analytic platform (280). At least one of these components of the intelligent server (200) may be omitted. Alternatively, the intelligent server (200) may additionally include other components.

[0061] The front end (215) can receive user input from the electronic device (201). The front end (215) can transmit a response corresponding to the user input.

[0062] The natural language platform (220) may include an automatic speech recognition module (ASR module) (221), a natural language understanding module (NLU module) (223), a planner module (225), a natural language generator module (NLG module) (227), and a text-to-speech module (TTS module) (229). At least one of these components of the natural language platform (220) may be omitted. Alternatively, the natural language platform (220) may additionally include other components.

[0063] In one embodiment, the automatic speech recognition module (221) can convert input voice data received from the electronic device (201) into input text data. The natural language understanding module (223) can determine the user's intent based on the input text data corresponding to the input voice data. For example, the natural language understanding module (223) can determine the user's intent by performing syntactic analysis or semantic analysis on the input text data. The natural language understanding module (223) can determine the meaning of words extracted from the input text data corresponding to the user input by using linguistic features (e.g., grammatical elements) of morphemes or phrases, and can determine the user's intent by matching the meaning of the determined words to the intent. In this way, the natural language understanding module (223) can determine intent information corresponding to the user's utterance. The intent information can be information indicating the user's intent determined by interpreting the text. Intention information may include information indicating an action (or function) that a user wishes to execute using a device (e.g., electronic device (210)).

[0064] The planner module (225) can generate a plan based on the intent information determined by the natural language understanding module (223). Based on the determined intent information, the planner module (225) can determine a plurality of operations included in each of a plurality of domains required to perform a task. The planner module (225) can determine parameters required to execute the determined plurality of operations and / or result values ​​output by the execution of the plurality of operations. The parameters and result values ​​can be defined as concepts of a specified format (or class). Accordingly, the plan can include a plurality of operations and a plurality of concepts determined by the user's intent. The planner module (225) can determine the relationship between the plurality of operations and the plurality of concepts in a step-by-step (or hierarchical) manner. For example, the planner module (225) can determine the execution order of the plurality of operations determined based on the user's intent based on the plurality of concepts. The planner module (225) can determine the execution order of multiple operations based on the parameters required for executing the multiple operations and the results output by executing the multiple operations. Accordingly, the planner module (225) can generate a plan including association information (e.g., ontology) between the multiple operations and the multiple concepts. The planner module (225) can generate the plan using information stored in a capsule database (230) in which a set of relationships between concepts and operations is stored.

[0065] The natural language generation module (227) can convert specified information into text format. The information converted into text format may be in the form of natural language speech. The text-to-speech conversion module (229) can convert text data into speech data.

[0066] In one embodiment, some or all of the functions of the natural language platform (220) described above may also be implemented in the electronic device (201). For example, the electronic device (201) may operate in the form of an on-device and perform the functions of the natural language platform (220) on its own without being connected to an intelligent server (200).

[0067] The capsule database (230) may store one or more capsules containing information about relationships between multiple concepts and actions corresponding to multiple domains. A capsule may include multiple action objects (or action information) and concept objects (or concept information) included in a plan. In one embodiment, the capsule database (230) may store multiple capsules in the form of a concept action network (CAN). The multiple capsules may be stored in a function registry included in the capsule database (230).

[0068] The capsule database (230) may include a strategy registry that stores strategy information necessary for determining a plan corresponding to user input, such as input voice data. The strategy information may include reference information for determining a single plan when there are multiple plans corresponding to the user input. In one embodiment, the capsule database (230) may include a follow-up registry that stores information on follow-up actions for suggesting follow-up actions to a user in a given situation. The follow-up actions may include, for example, follow-up utterances. In one embodiment, the capsule database (230) may include a vocabulary registry that stores vocabulary information included in the capsule information and / or a dialog registry that stores information on dialogue (or interaction) with the user. In one embodiment, the capsule database (230) may also be implemented in the electronic device (201).

[0069] The execution engine (240) can use the generated plan to produce results. The end user interface (250) can transmit the produced results to the electronic device (201). The electronic device (201) can receive the corresponding results transmitted from the intelligent server (200) and provide the received results to the user.

[0070] The management platform (260) can manage information used in the intelligent server (200). The big data platform (270) can collect user data. The analysis platform (280) can manage the quality of service (QoS) of the intelligent server (200). For example, the analysis platform (280) can manage the components and processing speed (or efficiency) of the intelligent server (200).

[0071] The service server (300) can provide services (e.g., food ordering or hotel reservation) specified for the electronic device (201). Services of the service server (300), such as CP Service A (301) and CP Service B (302), can interact with the front end (215) of the intelligent server (200).

[0072] In the integrated intelligent system (20) described above, the electronic device (201) can provide various intelligent services to the user in response to user input. The user input may include, for example, input via a physical button, touch input, or voice input.

[0073] In one embodiment, the electronic device (201) can provide a voice recognition service through an intelligent app (or voice recognition app) stored therein. For example, the electronic device (201) can recognize a user utterance or voice input received through a microphone (206) and provide a service corresponding to the recognized voice input to the user. The electronic device (201) can perform a designated operation based on the received voice input, either alone or together with the intelligent server (200) and / or the service server (300). For example, the electronic device (201) can execute an application corresponding to the received voice input and perform a designated operation through the executed application.

[0074] In one embodiment, the electronic device (201) may provide a translation service through an intelligent app (or translation app) stored within the device. The translation service may include, for example, a voice translation service and a real-time interpretation service. A real-time interpretation service is a service that facilitates communication between multiple languages ​​by interpreting phone calls or text messages in real time. Among the real-time interpretation services, a translation service during a call is a service that enables smooth communication between users who are making a call in different languages ​​through voice recognition and translation functions. For example, a translation service during a call may be a service that translates the language used by the other party in a call into the user's own language in real time when the user makes or receives a call.

[0075] In one embodiment, the processes for the translation service provided by the electronic device (201) can be performed independently by the electronic device (201) without connection to the intelligent server (200). In FIG. 2b below, an example is provided in which the electronic device (201) provides the translation service independently without the assistance of another device (e.g., the intelligent server (200)) in the form of an on-device. However, the scope of the embodiment is not limited thereto, and the electronic device (201) may also provide the translation service in collaboration with the intelligent server (200).

[0076] FIG. 2b is a block diagram illustrating a configuration of an electronic device that provides a translation service according to various embodiments.

[0077] In one embodiment, each user's electronic device (e.g., a smartphone) can recognize speech for a conversation (e.g., a call), translate the content of the recognized speech into the language of the conversation partner or a language selected by the user, and then transmit the translation result to the electronic device of the conversation partner. In one embodiment, the electronic device (210) can obtain the user's input voice data through the microphone (206) and convert the input voice data into input text data using STT (speech to text) technology. The electronic device (210) can translate the input text data into the language of the conversation partner or a language selected by the user to generate the translation result text data. For example, the translation result text data can be output to the screen through the display of the display module (204). Alternatively, the electronic device (210) can convert the translation result text data into voice output data using TTS (text to speech) technology. The electronic device (210) can transmit the translation result text data and / or the voice output data to the electronic device of the conversation partner. The other party's electronic device can display the received translation result text data on the screen, or convert the translation result text data into voice output data and output the voice output data through a speaker.

[0078] During the translation service, input voice data may be converted into input text data through voice recognition and stored. Translation is performed on the stored input text data. However, in a noisy environment (or an environment with a poor signal-to-noise ratio (SNR)) where sounds other than voice exist as noise when the voice is acquired, or in a situation where the user's pronunciation is unclear, the input text data may be converted differently from the user's speech, and in this case, the voice recognition rate and / or the accuracy of the translation result may decrease (e.g., the translation may be made into words that do not fit the context of the conversation). According to one or more embodiments described below, during the translation service, the electronic device (201) can distinguish whether the environment in which the voice is acquired is a noisy environment with a poor SNR. If it is determined to be a noisy environment, the electronic device (201) generates modified text data by modifying the input text data through a text modification model, and the translated text data may be generated based on the modified text data. The text correction model may be a model trained on user data (e.g., frequently used sentences, words, and / or speech patterns in previous conversations) and may be a model that modifies input text data by considering the current conversation context. User data, such as sentences, words, and / or speech patterns used by the user in previous conversations, may be recorded in a database, and the text correction model may be a model that has learned frequently used sentences, words, and / or speech patterns. The text correction model may output modified text data by modifying the input text data to make sense within the current conversation context by considering the current conversation context. The text correction model may be, for example, but is not limited to, a neural network model.By modifying input text data by considering user data and / or conversation context, smooth communication can be enabled and translation accuracy can be improved even in noisy environments with poor SNR or environments where the user's speech is not clear.

[0079] Additionally, in one embodiment, the electronic device (201) can improve the accuracy of the translation by comparing the translation result of the input text data with the translation result of the modified text data output from the text modification model based on the noise level (e.g., SNR) of the sound data input through the microphone (206), and selecting the translation result that is more suitable for the conversation context among the two translation results based on the comparison result.

[0080] In one embodiment, the electronic device (201) may include a microphone (206) for acquiring input voice data, a memory (207) including one or more storage media for storing instructions, and one or more processors (203). The one or more processors (203) may include processing circuitry. When the instructions are individually or collectively executed by the one or more processors (203), the instructions may cause the electronic device (201) or the one or more processors (203) to perform various operations.

[0081] For example, one or more processors (203) can convert input voice data of a first language acquired through a microphone (206) into input text data of the first language corresponding to the input voice data. One or more processors (203) can determine whether the surrounding environment when acquiring the input voice data is a noisy environment. To this end, one or more processors (203) can determine whether the noise level of the sound data input through the microphone (206) satisfies a condition. In one embodiment, the noise level of the sound data may refer to a noise level monitored in real time or may refer to the noise level of the sound data including the input voice data. The case where the noise level of the sound data satisfies the condition may include, for example, a case where the noise level (e.g., SNR) of the sound data is greater than a threshold value. When the noise level of the sound data is greater than the threshold value, it may be recognized that the surrounding environment is a noisy environment.

[0082] In one embodiment, if the noise level of the sound data does not satisfy the condition, one or more processors (203) may generate the translated text data in the second language by translating the input text data in the first language into the second language. The fact that the noise level does not satisfy the condition may indicate that the surrounding environment is not very noisy, and in this case, the translation result for the corresponding input text data may be output without modifying the input text data.

[0083] In one embodiment, when the noise level of the sound data satisfies the condition, one or more processors (203) may generate modified text data in a first language in which at least a portion of the text in the input text data has been changed using a neural network-based text modification model. In one embodiment, the text modification model may be trained using voice data collected for a speaker (e.g., a user of the electronic device (201)) corresponding to the input voice data as training data. The collected voice data may include, for example, information about sentences, words, and / or the speaker's tone uttered by the speaker in a previous conversation.

[0084] In one embodiment, the input of the text modification model may include input text data, and the output of the text modification model may include modified text data. If a previous conversation history exists in the current conversation, the input of the text modification model may further include conversation history data between the speaker and the speaker's conversation partner. The conversation history data may be text data that stores the conversation content from the beginning of the conversation to the present.

[0085] One or more processors (203) may generate translated text data in a second language based on modified text data output from a text modification model. For example, if the first language is Korean, the second language may be, but is not limited to, Korean, English, Japanese, or Chinese.

[0086] In one embodiment, one or more processors (203) may generate translated text data in a second language by translating the modified text data into a second language. One or more processors (203) may use a translator (e.g., the translator (577) of FIG. 5) to translate the modified text data in a first language into a second language, thereby generating translated text data in the second language.

[0087] In one embodiment, one or more processors (203) may generate translation result text data based on a translation result determined to be more appropriate between a translation result of input text data and a translation result of modified text data. The one or more processors (203) may translate input text data in a first language into a second language to generate first candidate translation result text data, and may translate the modified text data into the second language to generate second candidate translation result text data. A translator (e.g., the translator (577) of FIG. 5) may be used for translation into the second language. The one or more processors (203) may select either the first candidate translation result text data or the second candidate translation result text data as the translation result text data of the second language. For example, the one or more processors (203) may determine an evaluation index whose value is determined based on a difference between the first candidate translation result text data and the second candidate translation result text data, and may select the translation result text data of the second language based on the evaluation index. One or more processors (203) may determine an evaluation index based on the number of words or syllables that differ between the first candidate translation result text data and the second candidate translation result text data.

[0088] In one embodiment, one or more processors (203) may regard the second candidate translation result text data as the correct text, determine the number Dw of words incorrectly deleted from the first candidate translation result text, the number Sw of words incorrectly replaced, and the number Iw of words incorrectly added based on the second candidate translation result text data, and determine an evaluation index Ew according to the following mathematical expression 1 based on the number Nw of words in the second candidate translation result text data.

[0089]

[0090] Here, the evaluation index Ew can correspond to the word error rate (WER).

[0091] In one embodiment, one or more processors (203) may regard the second candidate translation result text data as the correct text, determine the number Dc of syllables incorrectly deleted, the number Sc of syllables incorrectly replaced, and the number Ic of syllables incorrectly added in the first candidate translation result text based on the second candidate translation result text data, and determine an evaluation index Ec according to the following mathematical expression 2 based on the number Nc of syllables in the second candidate translation result text data.

[0092]

[0093] Here, the evaluation index Ec can correspond to the character error rate (CER).

[0094] In one embodiment, when the evaluation index is less than or equal to a threshold value, one or more processors (203) may select the first candidate translation result text data as the translation result text data of the second language. When the evaluation index is greater than the threshold value, one or more processors (203) may select the second candidate translation result text data as the translation result text data of the second language. An evaluation index less than or equal to the threshold value may indicate that the difference between the first candidate translation result text data and the second candidate translation result text data is relatively small, and an evaluation index greater than the threshold value may indicate that the difference between the first candidate translation result text data and the second candidate translation result text data is relatively large. When the evaluation index is less than or equal to the threshold value, a translation result for input text data corresponding to the original text may be output, and when the evaluation index is greater than the threshold value, a translation result for text data modified by the text modification model may be output.

[0095] In one embodiment, the electronic device (201) may further include a speaker (205) that outputs voice output data. One or more processors (203) may convert the translation result text data into voice output data corresponding to the translation result text data, and control the speaker (205) to output the voice output data.

[0096] In one embodiment, the electronic device (201) may provide a translation service in conjunction with an intelligent server (200). When the electronic device (201) obtains input voice data of a first language through a microphone (206), the electronic device (201) may transmit the obtained input voice data to the intelligent server (200). The intelligent server (200) may convert the input voice data into input text data using an automatic speech recognition module (221). In one embodiment, the intelligent server (200) may measure a noise level of the input voice data or receive information about the noise level of the sound data measured by the electronic device (201) from the electronic device (210). If the noise level satisfies a condition (e.g., if the noise level is greater than a threshold value), the intelligent server (200) may generate modified text data of the first language in which at least a portion of the text in the input text data is changed using a text modification model. The intelligent server (200) may generate translated text data of the second language based on the modified text data. The process by which the intelligent server (200) generates the translation result text data may be the same as the process by which the electronic device (201) described above generates the translation result text data. The intelligent server (200) transmits the translation result text data to the electronic device (201), and the electronic device (201) may convert the translation result text data received from the intelligent server (200) into voice output data corresponding to the translation result text data. The electronic device (201) may output the voice output data through a speaker (205). The translation result text data may be converted into voice through voice synthesis technology and transmitted to the user.

[0097] FIG. 3 is a diagram for explaining outputting voice output data according to text-to-speech conversion during translation during a call according to various embodiments.

[0098] Referring to FIG. 3, an electronic device (310) according to one embodiment (e.g., the electronic device (101) of FIG. 1 or the electronic device (201) of FIG. 2B) may provide a translation function (or interpretation function) during a call with another electronic device (330). The electronic device (310) may translate in real time input voice data of a first user of the electronic device (310) and / or received voice data of a second user of the other electronic device (330) during a call.

[0099] In one embodiment, the electronic device (310) may display text data as a result of translation of input voice data of a first user on a display (e.g., a screen) (320) of the electronic device (310), and convert the text data as a result of translation of the first user into voice output data according to a text-to-speech (TTS) process. The electronic device (310) may provide the voice output data to another electronic device (330) in the form of a synthesized voice. In addition, the electronic device (310) may display text data as a result of translation of the voice data of a second user on the display (320) of the electronic device (310), and convert the text data as a result of translation of the voice data of the second user into voice output data. The electronic device (310) may provide the voice output data to the first user in the form of a synthesized voice.

[0100] In one embodiment, when translating input voice data (or voice signal) according to a first user's speech during a call in real time, the electronic device (310) may pause the speech at a certain point, convert the translated text data obtained by translating the input voice data up to the pause point into voice output data, and output the voice output data in the form of a synthesized voice. For convenience of explanation, it is assumed that the first user of the electronic device (310) utters in a first language, "I'm going to invite Jane to my birthday party on Friday. Can you give me Jane's contact information?" The electronic device (310) may display (322, 324, 326) the translated text data obtained by translating the input voice data according to the first user's speech of the electronic device (310) in real time into a second language, along with the result of voice recognition, on the display (320). The electronic device (310) can convert text data (e.g., "I'm going to invite Jane at a birthday party on Friday evening") translated in real time into a second language into voice output data and output it as a synthesized voice, from "I'm going to invite Jane at a birthday party on Friday evening. Can you give me Jane's contact information?"

[0101] In one embodiment, the electronic device (310) may convert the translation result text data into voice output data, and display an indicator (e.g., a user interface) (328) on the display (320) to control the output of a synthesized voice corresponding to the voice output data at a time when the voice output data is to be output. The electronic device (310) may mix the voice signal of “I’m going to invite Jane to a birthday party on Friday” with the synthesized voice of “I’m going to invite to Jane at a birthday party on Friday evening” and transmit the mixed signal to another electronic device (330). After outputting the voice output data for “I’m going to invite to Jane at a birthday party on Friday evening,” the electronic device (310) may also process the voice input data of “Can you give me Jane’s contact information?” in the same manner as described above.

[0102] FIG. 4 is a diagram for explaining an electronic device that outputs voice output data according to text-to-speech conversion during translation during a call according to various embodiments.

[0103] Referring to FIG. 4, an electronic device (410) according to one embodiment (e.g., the electronic device (101) of FIG. 1) may convert translation result text data into voice output data (or voice signal) and output it during translation (e.g., interpretation) during a call with another electronic device (e.g., the other electronic device (330) of FIG. 3). At this time, the electronic device (410) may determine the output of text-to-speech (TTS) for the translation result text data (e.g., the output timing of text-to-speech conversion).

[0104] According to one embodiment, the electronic device (410) may include a processor (420) (e.g., processor (120) of FIG. 1 or processor (203) of FIG. 2B), a memory (460) (e.g., memory (130) of FIG. 1 or memory (207) of FIG. 2B), an input module (412) (e.g., input module (150) of FIG. 1 or microphone (206) of FIG. 2B), an audio output module (414) (e.g., audio output module (155) of FIG. 1 or speaker (205) of FIG. 2B), a display module (450) (e.g., display module (160) of FIG. 1 or display module (204) of FIG. 2B), and / or an antenna module (416) (e.g., antenna module (197) of FIG. 1). For example, the first signal processing module (432), the second signal processing module (436), the Tx mixer (434), the Rx mixer (438), and the translation service (440) may be implemented as one or more of program code, an application, an algorithm, a routine, a set of instructions, or an artificial intelligence learning model, which are executable by the processor (420) and include instructions storable in the memory (460). For example, one or more of the first signal processing module (432), the second signal processing module (436), the Tx mixer (434), the Rx mixer (438), and the translation service (440) may be implemented as hardware and / or a combination of hardware and software.

[0105] According to one embodiment, the electronic device (410) may perform a transmission / reception process (or a speech process) for a call with another electronic device (e.g., another electronic device (330) of FIG. 3). The transmission / reception process may include a Tx process (or a speech process) that processes input voice data (or a voice signal) according to a speech of a first user of the electronic device (410) input by an input module (412) (e.g., a microphone), and an Rx process (or a speech process) that receives and processes a voice signal according to a speech of a second user of the other electronic device.

[0106] According to one embodiment, the Tx process may process input voice data through a first signal processing module (432), a translation service (440), and a Tx mixer (434). The input module (412) may receive input voice data according to speech in a first language. The first signal processing module (432) may perform signal processing on the input voice data received from the input module (412). For example, the first signal processing module (432) may perform signal processing on the input voice data using at least one of microphone array processing (MAP), adaptive echo canceller (AEC), noise suppression (NS), or automatic gain control or adaptive gain control (AGC). The translation service (440) may receive the input voice data processed by the first signal processing module (432) in real time, convert the received input voice data into input text data in a first language, and translate the input text data in the first language into a second language.

[0107] In one embodiment, when generating translation result text data of a second language, the translation service (440) may measure the noise level of sound data input through the input module (412), and if the noise level satisfies a condition, may generate modified text data of a first language in which at least a portion of the text in the input text data has been changed using a neural network-based text modification model. The translation service (440) may generate translation result text data of a second language based on the modified text data. In one embodiment, the translation service (440) may translate the modified text data of the first language into the second language to generate the translation result text data. Alternatively, the translation service (440) may translate input text data of the first language into the second language to generate first candidate translation result text data, translate the modified text data into the second language to generate second candidate translation result text data, and then select either the first candidate translation result text data or the second candidate translation result text data as the translation result text data of the second language. The process of determining or selecting the second translation result text data is described in more detail in FIGS. 7 to 9.

[0108] In one embodiment, the translation service (440) can convert the second language translation result text data, which is the result of translation into the second language, into voice output data (or voice output signal) in the second language. The Tx mixer (434) can mix the input voice data processed from the first signal processing module (432) and the second language voice output data at a set ratio to generate one output audio signal. The output audio signal can be transmitted to another electronic device (e.g., another electronic device (330) of FIG. 3) that is in a call with the electronic device (310) via the antenna module (416).

[0109] In one embodiment, the result processed by the translation service (440) of the Tx process may be displayed on the screen of the electronic device (410) by the display module (450). For example, the result of input voice data processed by the first signal processing module (432) being converted into input text data of a first language in real time and the result of input text data of the first language being translated into a second language in real time to determine the translation result text data of the second language may be displayed on the screen of the electronic device (410).

[0110] In the Rx process, voice data (or voice signal) received from another electronic device (e.g., another electronic device (330) of FIG. 3) may be processed through the second signal processing module (436), the translation service (440), and / or the Rx mixer (438). The antenna module (416) may receive voice data according to speech in a second language from another electronic device that is in a call with the electronic device (310). The second signal processing module (436) may perform signal processing on the voice data received from the antenna module (416). For example, the second signal processing module (436) may perform signal processing on the voice data using at least one of noise suppression (NS) and automatic gain control or adaptive gain control (AGC). The translation service (440) may convert voice data processed by the second signal processing module (436) into text data of a second language in real time, and translate the text data of the second language into a first language to generate text data of the translation result of the first language. In one embodiment, the translation service (440) may convert the translated text data of the first language into voice output data of the first language using a text-to-speech (TTS) process. The Rx mixer (438) may mix the voice data processed from the second signal processing module (436) and the voice output data of the first language at a set ratio to generate one output audio signal. The output audio signal may be output to the first user of the electronic device (410) through the audio output module (414) (e.g., a speaker).

[0111] In one embodiment, the result processed by the translation service (440) of the Rx process may be displayed on the screen of the electronic device (410) by the display module (450). For example, the result of voice data processed by the second signal processing module (436) being converted into text data of a second language in real time and the result of the text data of the second language being translated into a first language in real time to determine the translated result text of the first language may be displayed on the screen of the electronic device (410).

[0112] FIG. 5 is a diagram for explaining a translation service according to various embodiments.

[0113] Referring to FIG. 5, a translation service (550) according to one embodiment (e.g., the translation service (440) of FIG. 4) may be used for an application (e.g., a translation App) for the translation service (550). In addition, the translation service (550) may be used for another application (530) that requires the translation service (550). The application (530) may include, for example, one or more applications (e.g., an application that can use the translation service, such as a Call App, a Message App, a Note App, a video conferencing App, a recording App, or a chat App) running on an electronic device (e.g., the electronic device (101) of FIG. 1 , the electronic device (201) of FIG. 2B , the electronic device (310) of FIG. 3 , or the electronic device (410) of FIG. 4 ). The application (530) may transmit and receive information using an application program interface (API) to use the translation service (550). For example, an application (530) can use a real-time translation service by calling an API.

[0114] In one embodiment, the first voice signal (510) and / or the second voice signal (520) may be processed via the translation service (550). The first voice signal (510) may be a signal that is processed by receiving an utterance in a first language spoken by a first user using an electronic device (e.g., the electronic device (410) of FIG. 4) by an input module (e.g., the input module (412) of FIG. 4) of the electronic device. For example, the first voice signal (510) may include a voice signal that is processed by a first signal processing module (432) in a Tx process. The second voice signal (520) may be a signal that is processed by receiving an utterance in a second language spoken by a second user of another electronic device (e.g., the other electronic device (330) of FIG. 3) on a call with the first user by receiving an utterance in a second language by an antenna module (e.g., the antenna module (416) of FIG. 4) of the electronic device. The second voice signal (520) may include a voice signal processed from the second signal processing module (436) in the Rx process.

[0115] According to one embodiment, the translation service (550) may include a language pack (560), a speech information extractor (571), an automatic speech recognition (ASR) module (572) (e.g., a first ASR module (573) and a second ASR module (575)), a corrector (576), a translator (577), a TTS output determiner (579), and / or a TTS module (580).

[0116] The language pack (560) can support languages ​​for the real-time translation service provided by the translation service (550). The user can select the languages ​​used by the first user (e.g., the call transmitter) and the second user (e.g., the call receiver) to use the real-time translation service during a call. Additionally, the user can select the languages ​​used based on the contacts and / or address book stored in the electronic device. The user can set the languages ​​used in the settings screen. For example, when Jane (e.g., an English-speaking user) receives a call and the user (e.g., a Korean-speaking user) wants to use the "real-time translation during a call" service, Jane's language can be set from English to Korean, and the user's language can be set from Korean to English. In this case, the language pack (560) can pre-store the supported languages, and if the supported languages ​​are not available, new ones can be downloaded from a server that provides language data.

[0117] A voice information extractor (571) can extract voice information from a voice signal (e.g., a first voice signal (510) or a second voice signal (520)). The voice information extractor (571) can receive a voice signal in real time, extract voice information from the voice signal, and transmit the extracted voice information to an ASR module (572). In addition, the voice information extractor (571) can transmit the extracted voice information to a TTS output determiner (579).

[0118] In one embodiment, the voice information extractor (571) can extract voice information from a voice signal through various methods such as voice activity detection (VAD) and / or end point detection (EPD). The voice information can include, for example, information about a voice segment, information about a pause segment, the start time of speech, information about the start time of speech (e.g., information about the start time of speech, information about the end time of speech), intonation information (e.g., information about the pitch and / or low pitch), the end time of ASR, or any combination thereof.

[0119] In one embodiment, the ASR module (572) may perform ASR (e.g., ASR decoding) on ​​a speech signal (e.g., the first speech signal (510) or the second speech signal (520)) using the extracted speech information. For example, the ASR module (572) may perform ASR decoding (e.g., first-pass, second-pass decoding) on ​​a speech signal from the start of speech to the occurrence of a pause period. In addition, the ASR module (572) may perform ASR decoding on a speech signal from a pause period to the occurrence of another pause period, or from a pause period to the detected end of speech.

[0120] In one embodiment, the ASR module (572) may be implemented using various algorithms, such as a hidden markov model (HMM), weighted finite-state transducers (WFST), artificial neural networks (ANN), or support vector machines (SVM). For example, the ASR module (572) may be implemented using a neural network. For example, an RNN, a long short-term memory (LSTM), or a transformer may be used as the neural network, but there is no limitation on the types of neural networks that may be used.

[0121] In one embodiment, the ASR module (572) may include one or more ASR modules. For example, there may be one or more ASR modules (572) depending on the number and / or language of users performing the call (e.g., call transmitters or call receivers). For example, the ASR module (572) may include a first ASR module (573) and a second ASR module (575). The first ASR module (573) may perform ASR on a first voice signal (510), and the second ASR module (575) may perform ASR on a second voice signal (520). The first ASR module (573) may support a first language of a first user (e.g., call transmitter), and the second ASR module (575) may support a second language of a second user (call receiver). A first ASR module (573) can perform ASR on a first speech signal (510) of a first language to convert the first speech signal (510) into text data of the first language. A second ASR module (575) can perform ASR on a second speech signal (520) of a second language to convert the second speech signal (520) into text data of the second language.

[0122] In one embodiment, the translator (577) may receive the result of performing ASR from the ASR module (572) and translate the result of performing ASR into a target language based on a language supported by the ASR module (572). The result of performing ASR may be text data output by completing ASR decoding in real time by the ASR module (572). The translator (577) may perform translation on the output text data, and the translation result may be displayed through a display of the electronic device.

[0123] In one embodiment, the corrector (576) may receive text data in a first language as a result of performing ASR from the first ASR module (573), and perform correction on the text in the first language based on the surrounding circumstances. For example, the corrector (576) may determine whether the noise level of sound data input through a microphone of an electronic device (e.g., microphone (206) of FIG. 2B) satisfies a condition, and if the noise level satisfies the condition, the corrector (576) may generate corrected text data in the first language in which at least a portion of the text in the input text data has been changed using a neural network-based text correction model. If the noise level is greater than a threshold value, the noise level may be determined to satisfy the condition. The corrector (576) may request the translator (577) to generate translated text data in a second language based on the corrected text data.

[0124] In one embodiment, the translator (577) may receive a result of performing ASR (e.g., text data in a first language) from the first ASR module (573) and translate the result of performing ASR into a second language (e.g., translated result text data in the second language). In addition, the translator (577) may receive a result of performing ASR (e.g., text data in a second language) from the second ASR module (575) and translate the result of performing ASR into the first language (e.g., translated result text data in the first language). When the translator (577) receives a translation request for modified text data in the first language from the corrector (576), the translator (577) may translate the modified text data in the first language into the second language to generate translated result text data in the second language.

[0125] In one embodiment, the TTS output determiner (579) can receive the result of performing ASR (e.g., text data converted while performing ASR) from the ASR module (572). The TTS output determiner (579) can receive voice information (e.g., information about a short pause section, a start time of speech, an end time of speech, or an ASR end time) in real time. The voice information can be received directly from the voice information extractor (571), extracted from the voice information extractor (571) and received through the ASR module (572), or received from the ASR module (572) having a built-in voice information extraction function (e.g., the voice information extractor (571)).

[0126] According to one embodiment, the TTS output determiner (579) may determine to output the translation result (e.g., the final translation result) up to the section where the ASR no longer changes by using the text data and voice information converted while performing ASR by converting the text data into speech (TTS). The TTS output determiner (579) may identify the end point of a sentence (e.g., a complete sentence) in the text data based on a pause section (e.g., a short pause section and / or an EPD) in the text data output while performing ASR. The TTS output determiner (579) may determine to output the text-to-speech conversion for the sentence based on the point in time at which the end point of the sentence in the text data is identified based on the pause section in the text converted while performing ASR. The TTS output determiner (579) may control the TTS module (580) to output the text-to-speech conversion for the sentence. The TTS output determiner (579) can control the TTS module (580) to convert text to speech and output the sentence at each point in time when the end point of the sentence in the text data is identified.

[0127] In one embodiment, the TTS output determiner (579) may transmit text information (e.g., predicted punctuation information and complete sentences) and / or control information to the TTS module (580). The control information may be for controlling the TTS module (580) to convert text to speech and output the sentence. The TTS module (580) may generate and output text data corresponding to the translated sentence from the translator (577) as voice output data in the form of synthesized sound according to the control of the TTS output determiner (579).

[0128] In one embodiment, the TTS module (580) receives translated text corresponding to a sentence from a translator (577) (e.g., translated result text data translated into a second language and / or translated result text data translated into a first language), and generates and outputs the received text data as a synthesized sound by passing it through a text analyzer (581), a prosody predictor (583), or a synthesizer (585). In the case of the Tx process, the synthesized sound can be transmitted to another electronic device that is on a call. In the case of the Rx process, the synthesized sound can be output to a user of the electronic device through an audio output module (e.g., a speaker of the audio output module (414) of FIG. 4).

[0129] FIG. 6 is a diagram illustrating an example of a method for determining the output of text-to-speech conversion during translation during a call according to various embodiments.

[0130] In FIG. 6, it is assumed that a Tx process is processed in which a user of an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2b, the electronic device (310) of FIG. 3, or the electronic device (410) of FIG. 4) utters "I'm going to invite Jane to my birthday party on Friday evening. Can you tell me Jane's contact information?" during a call. The voice signal of the user's utterance "I'm going to invite Jane to my birthday party on Friday evening. Can you tell me Jane's contact information?" may be processed by a signal processing module (e.g., the first signal processing module (432) of FIG. 4) and input to the first ASR module (573).

[0131] In one embodiment, the first ASR module (573) can perform automatic speech recognition (ASR) on a voice signal received in real time, and sequentially output partial text data "Friday evening" (611), "To the birthday party" (612), "Jane" (613), and "I will invite you" (614) as a result of the ASR to the translator (577) and the TTS output determiner (579) as soon as the ASR is completed. The first ASR module (573) can sequentially output "Jane's contact information" (615) and "Can you tell me?" (616) to the translator (577) and the TTS output determiner (579). Additionally, the first ASR module (573) may sequentially output the partial texts “Friday evening”, “Friday evening birthday party”, “Jane to the Friday evening birthday party”, and “I will invite Jane to the Friday evening birthday party” to the translator (577) and the TTS output determiner (579) as soon as the ASR is completed.

[0132] In one embodiment, the corrector (576) may receive partial text data as a result of ASR from the first ASR module (573) and determine whether to correct the partial text data based on the noise level of sound data input through a microphone of the electronic device (e.g., microphone (206) of FIG. 2B). For example, if the noise level is greater than a threshold, the corrector (576) may correct the partial text data using a neural network-based text correction model. The text correction model may correct the partial text data by considering, for example, the conversation context up to the present in an ongoing call and / or the user's speech characteristics. The corrector (576) may request the translator (577) to translate the corrected partial text data into a language set by the user.

[0133] The first ASR module (573) can transmit voice information along with the ASR result to the TTS output determiner (579). The voice information can be extracted by the voice information extractor (571) or the first ASR module (573) can extract the voice information during ASR decoding. The first ASR module (573) can transmit information on pause sections (e.g., [SP], [EOS]) along with text information to the TTS output determiner (579). For example, the first ASR module (573) can transmit information of a pause interval between “Friday evening” (611) and “Birthday party” (612) (e.g., [SP] (621)), information of a pause interval between “Birthday party” (612) and “Jane” (613) (e.g., [SP] (622)), information of a pause interval between “Jane” (613) and “I will invite” (614) (e.g., [SP] (623)), information of a pause interval between “I will invite” (614) and “Jane’s contact” (615) (e.g., [SP] (624)), information of a pause interval between “Jane’s contact” (615) and “Can you tell me?” (616) (e.g., [SP] (625)), information of a pause interval after “Can you tell me” (e.g., [EOS] (626)) to the TTS output determiner (579) together with text information.

[0134] The TTS output determiner (579) can determine the point at which the sentence no longer changes based on the text (e.g., partial text) and speech information (e.g., information on a pause section or intonation information) that is the ASR result. For example, if the EOS result is provided together with the ASR result, the TTS output determiner (579) can determine that the user's speech has ended and request the TTS module (580) to convert the translation of the ASR result into text-to-speech and output it. As another example, if short pause information (e.g., [SP]) is provided as information on a pause section in the ASR result, the TTS output determiner (579) can analyze the previous text data at the time when the SP is recorded to determine that it is a complete sentence. The TTS output determiner (579) can perform the operations of sentence separation (or segmentation) and punctuation mark prediction (or punctuation mark insertion). The TTS output determiner (579) can determine that the text is a complete sentence by analyzing the previous text from the time when the pause section (e.g., [SP](621)) is captured to the time when the pause section (e.g., [EOS]) is captured. The TTS output determiner (579) can determine that the previous text (e.g., "Friday evening" (611), "At the birthday party on Friday evening" (611, 612), "At the birthday party on Friday evening" (611-613)) at the time when the pause sections (e.g., [SP](621), [SP](622), [SP](623)) are captured is not a complete sentence. The TTS output determiner (579) determines that the punctuation mark of a period can be included at the point “I will invite” (614) in the previous text (e.g., “I will invite Jane to the birthday party on Friday evening” (611-614)) at the time when the pause section (e.g., [SP] (624)) is captured, and thus determines that “I will invite Jane to the birthday party on Friday evening” (611-614) is a complete sentence.If the TTS output decision unit (579) determines that the sentence entered so far is a complete sentence based on SP, it can request the TTS module (580) to convert the translation of the determined complete sentence into text-to-speech and output it.

[0135] Before the TTS output determiner (579) requests the TTS module (580) to output text-to-speech conversion, the display module (e.g., the display module (450) of FIG. 4) can output the text converted by the first ASR module (573) (e.g., "Friday evening" (611), "At a birthday party on Friday evening" (611, 612), "Jane at a birthday party on Friday evening" (611-613), "I'm going to invite Jane to a birthday party on Friday evening" (611-614)) and the text translated by the translator (577) (e.g., "Friday evening" (631), "At a birthday party on Friday evening" (632), "Jane at a birthday party on Friday evening" (633), "I'm going to invite to Jane at a birthday party on Friday evening" (634)) together on the display (640) in real time. The results displayed on the display (640) may include temporary results (e.g., "Friday night" (611), "Friday night birthday party" (611, 612), "Jane at the Friday night birthday party" (611-613)) and / or final results (e.g., "I will invite Jane to the Friday night birthday party" (611-614)) that continuously change as the first ASR module (573) continuously streams out the ASR results. The point in time when the TTS output determiner (579) requests the output of the text-to-speech conversion to the TTS module (580) may be the point in time when the ASR results and / or the translation results (e.g., the translation of the ASR results) no longer change. At the point in time when the TTS module (580) requests the output of the text-to-speech conversion, an indicator (e.g., indicator (328) of FIG. 3) (e.g., UI) regarding the output of the text-to-speech conversion may be generated and displayed on the display (640). The indicator is intended to control the output of the synthesized sound generated by text-to-speech conversion, and the user can control the output of the synthesized sound through the indicator.

[0136] FIGS. 7 and 8 are flowcharts illustrating operations of a method for providing a translation service performed by an electronic device according to various embodiments. At least one of the operations in FIGS. 7 and 8 may be performed concurrently or in parallel with another operation, and the order of the operations may be changed. Furthermore, at least one of the operations may be omitted, and another operation may be additionally performed.

[0137] Referring to FIG. 7, in operation (710), an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 2B, or the electronic device (410) of FIG. 4) may obtain input voice data of a first language through a microphone of the electronic device (e.g., the microphone (206) of FIG. 2B). In one embodiment, a user of the electronic device may activate a translation function (or interpretation function) of the electronic device when making a call to another user and speak. The electronic device may obtain input voice data corresponding to the user's voice signal through the microphone.

[0138] In operation (720), a processor of an electronic device (e.g., processor (120) of FIG. 1, processor (203) of FIG. 2B, or processor (420) of FIG. 4) may convert input speech data into input text data of a first language corresponding to the input speech data. The processor may convert the input speech data into input text data, for example, using automatic speech recognition (e.g., ASR module (572) of FIG. 5). The processor may store the input text data. The processor may convert the content of a call between a user and another user, including the input text data, into text data and store the converted text data as conversation history data. The stored conversation history data may be used when determining the translated text data by considering the conversation context.

[0139] In operation (730), the processor may determine whether the noise level of sound data input through the microphone of the electronic device satisfies a condition. The sound data for which the noise level is measured may include sound data at the time of acquiring the input voice data or sound data input through the microphone in real time. The processor may measure a noise level, such as a signal-to-noise ratio (SNR), for the sound data, for example. A case in which the noise level of the sound data satisfies the condition may include, for example, a case in which the noise level of the sound data is greater than a threshold value. The processor may determine whether the environment in which the user speaks is a noisy environment based on the noise level of the sound data.

[0140] If the noise level is determined to satisfy the condition (yes in operation (730)), in operation (740), the processor may generate modified text data in the first language in which at least a portion of the text in the input text data has been changed using a neural network-based text modification model. The text modification model may be trained using voice data collected for a speaker (e.g., a user of an electronic device) corresponding to the input voice data as training data. The text modification model may learn voice data of a user of an electronic device in normal times and perform voice recognition prediction based on the learned content.

[0141] In one embodiment, the input of the text correction model may include input text data, and the output of the text correction model may include modified text data. For example, the input of the text correction model may further include conversation history data between a speaker and the speaker's conversation partner (e.g., another user on a call with the user of the electronic device). The text correction model may be a model that has learned the speech tone, vocabulary, and / or intonation of the user of the electronic device through a learning process, and may output modified text data that has been modified based on the speech tone, vocabulary, and / or intonation of the learned user based on the input text data input to the text correction model. Furthermore, when the text correction model receives input text data and conversation history data as inputs, it may modify the input text data by considering not only the speech tone, vocabulary, and / or intonation of the user, but also the conversation history (or conversation context) up to this point.

[0142] In operation (750), the processor may generate translated text data in a second language based on the modified text data. In one embodiment, the processor may generate translated text data in the second language by translating the modified text data into the second language. The processor may perform translation of the modified text data using the translation service (440) described in FIG. 4 or the translator (577) of FIG. 5.

[0143] In another embodiment, the processor may select to output the final translation result between the translation result for the input text data and the translation result for the modified text data. This is described in more detail below with reference to FIG. 8.

[0144] Referring to FIG. 8, in operation (810), the processor may translate input text data in a first language into a second language to generate first candidate translation result text data. In operation (820), the processor may translate modified text data into the second language to generate second candidate translation result text data.

[0145] The processor may select either the first candidate translation result text data or the second candidate translation result text data as the translation result text data of the second language. In one embodiment, in operation (830), the processor may determine an evaluation index whose value is determined based on a difference between the first candidate translation result text data and the second candidate translation result text data, and select the translation result text data of the second language from among the first candidate translation result text data and the second candidate translation result text data based on the evaluation index. The processor may determine the evaluation index based on the number of words or syllables that differ between the first candidate translation result text data and the second candidate translation result text data, as described in FIG. 2B. The evaluation index may include, for example, a word error rate (WER) or a syllable error rate (CER).

[0146] In operation (840), the processor may determine whether the determined evaluation index is less than or equal to a threshold value. If the evaluation index is less than or equal to the threshold value ("Yes" in operation (840)), the processor may select the first candidate translation result text data as the translation result text data of the second language in operation (850). If the evaluation index is greater than the threshold value ("No" in operation (840)), the processor may update the conversation history data based on the modified text data generated in operation (740) in operation (860). For example, the processor may replace the input text data stored in the conversation history data with the modified text data. In operation (870), the processor may select the second candidate translation result text data as the translation result text data of the second language.

[0147] Returning to FIG. 7, if it is determined that the noise level does not satisfy the condition (if 'No' in operation (730)), in operation (760), the processor can generate the translation result text data in the second language based on the input text data. For example, if the noise level of the sound data is below a threshold value, it can be determined that the noise level does not satisfy the condition. The processor can generate the translation result text data in the second language by translating the input text data in the first language into the second language. The processor can perform the translation on the input text data using the translation service (440) described in FIG. 4 or the translator (577) of FIG. 5. If the noise level does not satisfy the condition, it can be determined that the environment in which the user speaks is not noisy, in which case the processor can translate the input text data as is without modifying the input text data to generate the translation result text data in the second language.

[0148] In operation (770), the processor may convert the translation result text data into voice output data corresponding to the translation result text data. For example, the processor may convert the translation result text data into voice output data using TTS technology (e.g., the TTS output determiner (579) of FIG. 5 or the TTS module (580)).

[0149] In operation (780), the processor may transmit voice output data to another electronic device, which is a device of a conversation partner, through a communication module (e.g., communication module (190) of FIG. 1) or interface (e.g., interface (202) of FIG. 2b) of the electronic device.

[0150] According to the above-described embodiments, the electronic device can improve the accuracy of the translation result by considering the user's speech characteristics (e.g., frequently used sentences, words, and speech patterns of the user) and / or the conversation context between the user and the conversation partner for sentence prediction, thereby enabling smooth communication between the user and the conversation partner. In addition, the electronic device can determine whether the environment in which the user speaks is a noisy environment that may lower the accuracy of speech recognition based on the noise level of sound data acquired through the microphone, and can improve the accuracy of the translation result by modifying the speech recognition result in a noisy environment based on the user's speech characteristics and / or the conversation context.

[0151] FIG. 9 is a flowchart illustrating operations of a method performed by an electronic device providing a translation service according to various embodiments. In one embodiment, at least one of the operations in FIG. 9 may be performed concurrently or in parallel with another operation, and the order of the operations may be changed. Furthermore, at least one of the operations may be omitted, and another operation may be additionally performed.

[0152] In operation (910), an electronic device (e.g., an electronic device (101) of FIG. 1, an electronic device (201) of FIG. 2B, or an electronic device (410) of FIG. 4) may perform a phone call with a counterpart device, which is a device of a user's conversation partner (or call partner), when a phone call is connected. The electronic device may obtain input voice data corresponding to a voice signal spoken by a user of the electronic device through an input module (e.g., an input module (150) of FIG. 1, or an input module (412) of FIG. 4) or a microphone (e.g., a microphone (206) of FIG. 2B). Alternatively, the electronic device may receive voice data of the conversation partner from the counterpart device through a communication module (e.g., a communication module (190) of FIG. 1) or an interface (e.g., an interface (202) of FIG. 2B).

[0153] In operation (915), the processor of the electronic device (e.g., the processor (120) of FIG. 1, the processor (203) of FIG. 2B, the processor (420) of FIG. 4) can determine whether the interpretation function for providing the translation service is on (or activated). If the interpretation function is off (or deactivated) rather than on ('No' in operation (915)), the processor can perform a phone call with the other party device without performing the translation service. In operation (990), the processor can check whether the call with the other party device is terminated, and if the call is not terminated, monitor whether the interpretation function is turned on according to operation (915).

[0154] If the interpretation function is turned on (if 'yes' in operation (915)), in operation (920), the processor can convert the conversation content into text data and store the text data. The processor can convert input voice data in a first language corresponding to a voice signal spoken by a user into input text data in the first language corresponding to the input voice data, and store the input text data. The processor can convert voice data corresponding to a voice signal of the conversation partner transmitted from the other party's device into text data, and store the text data. The processor can convert the conversation content during a phone call into the form of text data using a voice recognition function, and store the text data for all of the conversation content.

[0155] In operation (925), the processor may determine whether the current conversation data to be translated is conversation data transmitted from the other party's device. The processor may determine whether the current subject to be translated is input text data corresponding to the user's utterance or text data corresponding to the other party's utterance.

[0156] If the target of translation is conversation data (text data corresponding to the other party's speech) transmitted from the other party's device (if 'yes' in operation (925)), in operation (930), the processor can generate translation result text data in the second language by translating the conversation data into the second language. In operation (935), the processor can convert the translation result text data into voice output data using TTS technology. In operation (940), the processor can output the voice output data through an audio output module (e.g., audio output module (155) of FIG. 1, or audio output module (414) of FIG. 4) or a speaker (e.g., speaker (205) of FIG. 2b).

[0157] If the target to be translated is input text data corresponding to the user's speech (if 'No' in operation (925)), the processor may measure the noise level of the sound data in operation (945). The processor may, for example, monitor sound data heard through a microphone and measure the noise level, such as SNR, for the sound data.

[0158] In operation (950), the processor can determine whether the noise level of the sound data is greater than a threshold value.

[0159] If the noise level of the sound data is not greater than the threshold value (if 'No' in operation (950)), in operation (955), the processor can generate translated text data in the second language by translating the input text data into the second language. In operation (960), the processor can convert the translated text data into voice output data. In operation (965), the processor can transmit the voice output data to the counterpart device through a communication module (e.g., the communication module (190) of FIG. 1) or an interface (e.g., the interface (202) of FIG. 2B).

[0160] If the noise level of the sound data is greater than the threshold value (if 'Yes' in operation (950)), in operation (970), the processor may generate modified text data in the first language in which at least a portion of the text in the input text data has been changed using a text modification model. The text modification model may be a model that has learned the user's speech characteristics through user learning data (e.g., the user's speech data showing sentences, words, and / or speech patterns (e.g., dialect, tone, or intonation) frequently used by the user in normal conversations). The text modification model may generate the modified text data by modifying the input text data according to the conversation context and / or the user's speech characteristics.

[0161] Improving the accuracy of speech recognition for user-spoken electronic devices can lead to more accurate translation results. For example, for users with inaccurate pronunciation, a text correction model can learn the user's speech characteristics in advance and correct or supplement the speech recognition results accordingly, resulting in more accurate translation results. However, even if the user's speech characteristics are learned in advance, it can be difficult to clearly distinguish the user's voice in noisy environments, potentially resulting in poor speech recognition. To address these issues, a text correction model can additionally learn speech characteristics, such as voice frequency characteristics, specific to the user's pronunciation. By assigning weights to each pronunciation, the text correction model can obtain corrected text data with improved speech recognition accuracy. The text correction model can learn speech characteristics, such as accentuating the voice signal corresponding to specific frequencies when a user utters specific words and / or sentences. When analyzing input text data, the text correction model can generate corrected text data by assigning higher weights to specific frequencies based on previously learned user speech characteristics.

[0162] In operation (975), the processor may translate input text data in a first language into a second language to generate first candidate translation result text data, translate modified text data into the second language to generate second candidate translation result text data, and then determine the translation result text data in the second language based on the first candidate translation result text data and the second candidate translation result text data. The processor may determine an evaluation index whose value is determined based on a difference between the first candidate translation result text data and the second candidate translation result text data, and select the translation result text data from among the first candidate translation result text data and the second candidate translation result text data based on the evaluation index.

[0163] In one embodiment, the processor may determine an evaluation index, such as a word error rate and / or a syllable error rate, based on the number of words or syllables that differ between the first candidate translation result text data and the second candidate translation result text data. The processor may compare the evaluation index with a threshold value, and select the first candidate translation result text data as the translation result text data if the evaluation index is less than or equal to the threshold value, and select the second candidate translation result text data as the translation result text data if the evaluation index is greater than the threshold value. If the second candidate translation result text data is selected as the translation result text data, the processor may update the conversation history data based on the modified text data generated in operation (970). For example, the processor may replace input text data stored in the conversation history data with the modified text data.

[0164] In operation (980), the processor may convert the translation result text data into voice output data. In operation (985), the processor may transmit the voice output data to the other party device through a communication module or interface.

[0165] After the above operation (940), operation (965), or operation (985), the processor may determine whether the call has ended in operation (990). If the call is determined not to have ended (i.e., 'No' in operation (990)), the processor may perform the operation again from operation (915).

[0166] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1, electronic device (201) of FIG. 2B, electronic device (410) of FIG. 4) may include a microphone for acquiring input voice data (e.g., input module (150) of FIG. 1, microphone (206) of FIG. 2B, input module (412) of FIG. 4), a memory including one or more storage media for storing instructions (e.g., memory (130) of FIG. 1, memory (207) of FIG. 2B, memory (460) of FIG. 4), and one or more processors including processing circuits (e.g., processor (120) of FIG. 1, processor (203) of FIG. 2B, processor (420) of FIG. 4).

[0167] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to convert input voice data of a first language acquired through the microphone into input text data of a first language corresponding to the input voice data, and, when a noise level of sound data input through the microphone satisfies a condition, generate modified text data of the first language in which at least a part of the text in the input text data is changed using a neural network-based text modification model, and generate translated result text data of a second language based on the modified text data.

[0168] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to translate input text data in the first language into the second language to generate first candidate translation result text data, translate the modified text data into the second language to generate second candidate translation result text data, and select either the first candidate translation result text data or the second candidate translation result text data as the translation result text data in the second language.

[0169] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to determine an evaluation index, the value of which is determined based on a difference between the first candidate translation result text data and the second candidate translation result text data, and to select the translation result text data of the second language based on the evaluation index.

[0170] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to determine the evaluation index based on the number of words or syllables that differ between the first candidate translation result text data and the second candidate translation result text data.

[0171] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to select the first candidate translation result text data as the translation result text data of the second language when the evaluation index is less than or equal to a threshold value.

[0172] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to select the second candidate translation result text data as the translation result text data of the second language if the evaluation index is greater than the threshold value.

[0173] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to generate translated text data in the second language by translating input text data in the first language into the second language if the noise level does not satisfy the condition.

[0174] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to generate translated text data in the second language by translating the modified text data into the second language.

[0175] In one embodiment, the text correction model may be trained using voice data collected for a speaker corresponding to the input voice data as training data.

[0176] In one embodiment, the input of the text modification model may include the input text data, and the output of the text modification model may include the modified text data.

[0177] In one embodiment, the input of the text modification model may further include conversation history data between the speaker and the speaker's conversation partner.

[0178] In one embodiment, the case where the noise level of the sound data satisfies the condition may include a case where the noise level of the sound data is greater than a threshold value.

[0179] In one embodiment, when the instructions are individually or collectively executed by the one or more processors, the instructions may cause the electronic device to convert the translation result text data into voice output data corresponding to the translation result text data.

[0180] In one embodiment, a method performed by an electronic device (e.g., an electronic device (101) of FIG. 1, an electronic device (201) of FIG. 2B, an electronic device (410) of FIG. 4) may include: acquiring input voice data of a first language through a microphone of the electronic device (e.g., an input module (150) of FIG. 1, a microphone (206) of FIG. 2B, an input module (412) of FIG. 4); converting the input voice data into input text data of the first language corresponding to the input voice data; determining whether a noise level of sound data input through the microphone satisfies a condition; generating modified text data of the first language in which at least a portion of text is changed in the input text data using a neural network-based text modification model when the noise level is determined to satisfy the condition; and generating translated result text data of a second language based on the modified text data.

[0181] In one embodiment, the operation of generating the translation result text data of the second language may include an operation of translating the input text data of the first language into the second language to generate first candidate translation result text data, an operation of translating the modified text data into the second language to generate second candidate translation result text data, and an operation of selecting one of the first candidate translation result text data and the second candidate translation result text data as the translation result text data of the second language.

[0182] In one embodiment, the operation of selecting one of the first candidate translation result text data and the second candidate translation result text data as the translation result text data of the second language may include the operation of determining an evaluation index whose value is determined according to a difference between the first candidate translation result text data and the second candidate translation result text data, and the operation of selecting the translation result text data of the second language based on the evaluation index.

[0183] In one embodiment, the operation of selecting the translation result text data of the second language based on the evaluation index may include an operation of selecting the first candidate translation result text data as the translation result text data of the second language when the evaluation index is less than or equal to a threshold value, and an operation of selecting the second candidate translation result text data as the translation result text data of the second language when the evaluation index is greater than the threshold value.

[0184] In one embodiment, the method may further include an operation of generating translated result text data in the second language by translating the input text data in the first language into the second language if the noise level does not satisfy the condition.

[0185] In one embodiment, the operation of generating the translation result text data in the second language may include an operation of generating the translation result text data in the second language by translating the modified text data into the second language.

[0186] In one embodiment, the method may further include an operation of converting the translation result text data into voice output data corresponding to the translation result text data.

[0187] In one embodiment, a computer-readable recording medium storing one or more computer programs may include instructions for performing the method.

[0188]

[0189] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0190] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0191] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0192] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more commands stored in a storage medium (e.g., an internal memory (136) or an external memory (138) of FIG. 1, a memory (207) of FIG. 2B, a memory (460) of FIG. 4)) readable by a machine (e.g., an electronic device (101) of FIG. 1, an electronic device (201) of FIG. 2B, an electronic device (410) of FIG. 4). For example, a processor (e.g., a processor (120) of FIG. 1, a processor (203) of FIG. 2B, a processor (420) of FIG. 4)) of a machine (e.g., an electronic device (101) of FIG. 1, an electronic device (201) of FIG. 2B, an electronic device (410) of FIG. 4)) may call at least one command among one or more commands stored from the storage medium and execute it. This enables the device to operate to perform at least one function according to at least one command called above. The one or more commands may include code generated by a compiler or code executable by an interpreter. The device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' only means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily in the storage medium.

[0193] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0194] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0195] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and software applications running on the operating system. Furthermore, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0196] Software may include computer programs, codes, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, or computer storage medium or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.

[0197] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.

[0198] The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiments, and vice versa.

[0199] While this disclosure has been illustrated and described with reference to various embodiments, it will be understood that the various embodiments are illustrative and not limiting. It will be further understood by those skilled in the art that various changes in form and detail may be made without departing from the true spirit and scope of the present disclosure, including the appended claims and their equivalents. Furthermore, it will be understood that any embodiment(s) described in this disclosure may be used in conjunction with any other embodiment(s) described in this disclosure.

Claims

1. In electronic devices (101; 201; 410), A microphone (150; 206; 412) for acquiring input voice data; A memory (130; 207; 460) comprising one or more storage media for storing instructions; and One or more processors (120; 203; 420) including processing circuitry Including, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: Converting the input voice data of the first language acquired through the above microphone (150; 206; 412) into input text data of the first language corresponding to the input voice data, If the noise level of the sound data input through the microphone (150; 206; 412) satisfies the condition, a text modification model based on a neural network is used to generate modified text data of a first language in which at least a portion of the text in the input text data has been changed, To generate text data of a second language translation result based on the above modified text data, Electronic devices (101; 201; 410).

2. In paragraph 1, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: Translating the input text data of the first language into the second language to generate first candidate translation result text data, Translating the above modified text data into the second language to generate second candidate translation result text data, Selecting either the first candidate translation result text data or the second candidate translation result text data as the translation result text data of the second language. Electronic devices (101; 201; 410).

3. In paragraph 2, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: Determine an evaluation index whose value is determined based on the difference between the first candidate translation result text data and the second candidate translation result text data, Selecting the translation result text data of the second language based on the above evaluation index, Electronic devices (101; 201; 410).

4. In paragraph 3, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: Determine the evaluation index based on the number of words or syllables that differ between the first candidate translation result text data and the second candidate translation result text data. Electronic devices (101; 201; 410).

5. In paragraph 3 or 4, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: If the above evaluation index is less than or equal to a threshold value, the first candidate translation result text data is selected as the translation result text data of the second language. Electronic devices (101; 201; 410).

6. In any one of paragraphs 3 to 5, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: If the above evaluation index is greater than the threshold value, the second candidate translation result text data is selected as the translation result text data of the second language. Electronic devices (101; 201; 410).

7. In any one of paragraphs 1 to 6, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: If the noise level does not satisfy the condition, the translation result text data of the second language is generated by translating the input text data of the first language into the second language. Electronic devices (101; 201; 410).

8. In any one of paragraphs 1 to 7, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: Generating the translated text data of the second language by translating the modified text data into the second language, Electronic devices (101; 201; 410).

9. In any one of paragraphs 1 to 8, The above text modification model is, It is learned by using the voice data collected for the speaker corresponding to the above input voice data as learning data, The input of the above text modification model includes the above input text data, The output of the above text modification model includes the modified text data. Electronic devices (101; 201; 410).

10. In paragraph 9, The input of the above text modification model is: Further including conversation history data between the speaker and the speaker's conversation partner, Electronic devices (101; 201; 410).

11. In any one of paragraphs 1 to 10, If the noise level of the above sound data satisfies the above conditions, Including cases where the noise level of the above sound data is greater than the threshold value, Electronic devices (101; 201; 410).

12. In any one of paragraphs 1 to 11, When the above instructions are individually or collectively executed by the one or more processors (120; 203; 420), the instructions cause the electronic device (101; 201; 410) to: Converting the above translation result text data into voice output data corresponding to the above translation result text data, Electronic devices (101; 201; 410).

13. In a method performed by an electronic device (101; 201; 410), An operation (710) of acquiring input voice data of a first language through a microphone (150; 206; 412) of the electronic device (101; 201; 410); An operation (720) of converting the input voice data into input text data of a first language corresponding to the input voice data; An operation (730) of determining whether the noise level of sound data input through the above microphone (150; 206; 412) satisfies a condition; If it is determined that the noise level satisfies the condition, an operation (740) of generating modified text data of a first language in which at least some text in the input text data has been changed using a neural network-based text modification model; and An operation (750) of generating a translation result text data of a second language based on the above modified text data. How to include.

14. In paragraph 13, The operation (750) of generating the translation result text data of the second language is as follows: An operation (810) of translating input text data of the first language into the second language to generate first candidate translation result text data; An operation (820) of translating the above modified text data into the second language to generate second candidate translation result text data; and An operation of selecting one of the first candidate translation result text data and the second candidate translation result text data as the translation result text data of the second language. How to include.

15. In paragraph 14, The operation of selecting one of the first candidate translation result text data and the second candidate translation result text data as the translation result text data of the second language is as follows: An operation of determining an evaluation index whose value is determined based on the difference between the first candidate translation result text data and the second candidate translation result text data; and Including an operation of selecting the translation result text data of the second language based on the evaluation index, The operation of selecting the translation result text data of the second language based on the above evaluation index is as follows: If the evaluation index is less than or equal to a threshold value, an operation of selecting the first candidate translation result text data as the translation result text data of the second language; and If the above evaluation index is greater than the above threshold value, an operation of selecting the second candidate translation result text data as the translation result text data of the second language. How to include.

Citation Information

Patent Citations

  • Methods, systems, and computer-readable recording media for cross-language image search options

    KR1020120135188A

  • Terminal and handsfree device for servicing handsfree automatic interpretation, and method thereof

    KR1020150026754A

  • Voice recognizing method and voice recognizing appratus

    KR1020160066441A

  • High performance sanitary pump

    KR102129695B1

  • KR20230160604A