Apparatus for translation using gestures, operation method and recording medium
The gesture-based translation system on smartwatches addresses the limitations of existing technologies by adapting translation to screen orientation and voice detection, providing efficient and context-aware language conversion.
Patent Information
- Application Number
- PCT/KR2025/006328
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-02
- Filing Date
- 2025-05-12
- Publication Date
- 2025-12-26
AI Technical Summary
Existing translation technologies on smartwatches lack efficient real-time translation capabilities that adapt to user interactions and context, such as screen orientation and voice detection, limiting their effectiveness in communication scenarios.
A gesture-based translation system that utilizes inertial sensors and microphones to determine screen direction and detect voice data, translating between languages based on user or counterparty orientation, and adjusting output accordingly.
Enables real-time, context-aware translation on smartwatches by accurately determining screen direction and voice data, enhancing communication efficiency and readability.
Smart Images

Figure KR2025006328_26122025_PF_FP_ABST
Abstract
Description
Translation device using gestures, its operating method and recording medium
[0001] The following embodiments relate to a translation device using gestures, an operating method thereof, and a recording medium.
[0002] As globalization accelerates in modern society, communication between people speaking different languages is becoming increasingly important. To meet this demand, interpretation and translation technology is rapidly advancing, and demand for real-time interpretation and translation services via mobile devices is particularly growing. Smartwatches are one such mobile device, offering users a variety of functions on their wrists.
[0003] Smartwatches feature a small display, microphone, speaker, and various sensors that can be worn on the wrist, allowing for easy access anytime, anywhere. This allows them to be used for a variety of purposes, including health management, exercise tracking, and notification reception. Recently, they are also evolving into offering interpretation and translation capabilities.
[0004] The above information is provided solely as background information to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0005] An aspect of the present disclosure is to address at least the problems and / or drawbacks mentioned above, and to provide at least some of the advantages described below. Accordingly, according to one aspect of the present disclosure, a gesture-based translation device, an operating method thereof, and a recording medium can be provided.
[0006] Additional aspects are presented in part in the description that follows, and in part will become apparent from the description or may be understood by practicing the embodiments provided.
[0007] According to one aspect of the present disclosure, a translation method using gestures is provided. The method may include: an operation of acquiring sensor data; an operation of determining a gesture of an electronic device and a state of the electronic device based on the sensor data; an operation of determining whether a screen direction of the electronic device is oriented toward a user or a counterparty based on the state of the electronic device when a gesture requesting translation is detected; an operation of determining whether acquired voice data exists; an operation of translating the acquired voice data into a user's language or a counterparty's language based on the screen direction of the electronic device if the acquired voice data exists; and an operation of outputting a translated text.
[0008] According to one aspect of the present disclosure, one or more computer programs stored on one or more non-transitory computer-readable storage media are provided. When the computer programs are individually or collectively executed by one or more processors of an electronic device, the computer programs cause the electronic device to perform the following operations. The operations may include: acquiring sensor data; determining a gesture of the electronic device and a state of the electronic device based on the sensor data; determining whether a screen direction of the electronic device is directed toward the user or toward the other party based on the state of the electronic device when a gesture requesting translation is detected; determining whether acquired voice data exists; translating the acquired voice data into the user's language or the other party's language according to the screen direction of the electronic device if the acquired voice data exists; and outputting a translated text.
[0009] According to one aspect of the present disclosure, an electronic device is provided. The electronic device includes an inertial sensor for measuring inertial sensor data; one or more microphones for receiving voice; and a memory for storing one or more computer programs; and one or more processors communicatively coupled to the inertial sensor, the one or more microphones, and the memory, wherein the one or more computer programs include computer-executable instructions, which, when individually or collectively executed by the one or more processors, cause the electronic device to perform the following operations: acquiring sensor data; determining a gesture of the electronic device and a state of the electronic device based on the sensor data; determining, when a gesture requesting translation is detected, whether a screen direction of the electronic device is directed toward a user or a counterparty based on the state of the electronic device; determining whether acquired voice data exists; translating, if acquired voice data exists, the acquired voice data into a language of the user or a language of the counterparty based on the screen direction of the electronic device; and outputting a translated text.
[0010] According to one aspect of the present disclosure, a translation method using a gesture according to one embodiment is provided. The method may include: acquiring sensor data of a wearable device; determining a gesture of the wearable device and a state of the wearable device through the sensor data; when a gesture requesting translation is detected by the wearable device, determining whether a screen direction of the wearable device is oriented toward a user or a counterparty through the state of the wearable device; determining whether voice data acquired by the wearable device exists; if the voice data acquired by the wearable device exists, transmitting the voice data acquired by the wearable device to a mobile device so as to translate the voice data into the user's language or the counterparty's language according to the screen direction of the wearable device; translating the voice data acquired by the mobile device from the wearable device into the user's language or the counterparty's language; and outputting a translated text by the mobile device.
[0011] Other aspects, advantages and important features of the present disclosure will become apparent to those skilled in the art from the following detailed description of various embodiments of the present disclosure taken in conjunction with the accompanying drawings.
[0012] The above-described side pins and other aspects, features and advantages of specific embodiments of the present disclosure will become more apparent from the following description taken in conjunction with the accompanying drawings.
[0013] FIG. 1 is a diagram illustrating the configuration of an electronic device according to an embodiment of the present disclosure.
[0014] FIG. 2 is a diagram illustrating the operation of a processor of an electronic device according to an embodiment of the present disclosure.
[0015] FIG. 3 is a diagram illustrating an operation of translating voice in a processor of an electronic device according to an embodiment of the present disclosure.
[0016] FIG. 4 is a flowchart illustrating an operation of performing translation according to gestures and screen orientation in an electronic device according to an embodiment of the present disclosure.
[0017] FIG. 5 is a flowchart illustrating an operation of performing translation according to gestures and screen orientation in an electronic device according to an embodiment of the present disclosure.
[0018] FIG. 6 is a flowchart illustrating an operation of checking a screen orientation in an electronic device according to an embodiment of the present disclosure.
[0019] FIG. 7 is a flowchart illustrating an operation of detecting a speaker's speech in an electronic device according to an embodiment of the present disclosure.
[0020] FIG. 8 is a flowchart illustrating an operation of translating a voice signal in an electronic device according to an embodiment of the present disclosure.
[0021] FIG. 9 is a diagram illustrating an example of determining a user direction and an opponent direction based on a sensor measurement value in an electronic device according to an embodiment of the present disclosure.
[0022] FIG. 10 is a drawing illustrating an example of providing a user's voice translated into the other party's language according to an embodiment of the present disclosure.
[0023] FIG. 11 is a drawing illustrating an example of providing a translation of the other party's voice into the user's language according to an embodiment of the present disclosure.
[0024] FIG. 12 is a drawing illustrating an example of outputting in a direction that can be read by a person viewing the screen according to the screen orientation according to an embodiment of the present disclosure.
[0025] FIG. 13 is a drawing illustrating an example of adjusting the font size according to the distance between an electronic device and a counterpart according to an embodiment of the present disclosure.
[0026] FIG. 14 is a drawing illustrating an example of sliding and outputting letters according to the distance between an electronic device and a counterpart according to an embodiment of the present disclosure.
[0027] FIG. 15 is a diagram illustrating an example of receiving a user's voice when the other party is present, according to an embodiment of the present disclosure.
[0028] FIG. 16 is a diagram illustrating an example of providing a translation of a user's voice into the language of a counterpart when the counterpart is present, according to an embodiment of the present disclosure.
[0029] FIG. 17 is a block diagram schematically illustrating the structure of an electronic device according to an embodiment of the present disclosure.
[0030] FIG. 18 is a block diagram illustrating an electronic device within a network environment according to an embodiment of the present disclosure.
[0031] FIG. 19 is a front perspective view of an electronic device according to an embodiment of the present disclosure.
[0032] FIG. 20 is a rear perspective view of an electronic device according to an embodiment of the present disclosure.
[0033] FIG. 21 is an exploded perspective view of an electronic device according to an embodiment of the present disclosure.
[0034] FIG. 22 is a diagram illustrating an example of a combination that provides a translation service through linkage between a wearable device that detects gestures and a mobile device according to an embodiment of the present disclosure.
[0035] FIG. 23 is a diagram illustrating an example of outputting a translated result according to an embodiment of the present disclosure.
[0036] FIG. 24 is a diagram illustrating an example of a gesture requesting translation in the case of a smart ring according to an embodiment of the present disclosure.
[0037] It should be noted that the same reference numbers are used throughout the drawings to indicate identical or similar elements, features, and structures.
[0038] The following description, with reference to the attached drawings, is intended to provide a comprehensive understanding of various embodiments of the present disclosure as defined by the claims and their equivalents. While this description includes various specific details to aid understanding, these should be considered merely exemplary. Accordingly, those skilled in the art will appreciate that various changes and modifications can be made to the various embodiments described herein without departing from the spirit and scope of the present disclosure. Furthermore, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0039] The terms and words used in the following description and claims are not limited to their dictionary meanings, but rather have been used by the inventors to clearly and consistently describe the technical content of the present disclosure. Therefore, those skilled in the art will readily understand that the following description of various embodiments of the present disclosure is intended solely for illustrative purposes and is not intended to limit the scope of the present disclosure, which is defined by the appended claims and their equivalents.
[0040] Additionally, unless the context clearly dictates otherwise, expressions in the singular should be interpreted to include the plural. For example, the expression "component surface" is interpreted to include one or more component surfaces.
[0041] The terminology used in this disclosure is for the purpose of description only and should not be construed as limiting. The singular expression includes the plural expression unless the context clearly indicates otherwise. In this specification, the terms "comprises" or "has" and the like are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood to not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0042] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art to which the embodiments pertain. Terms defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0043] In addition, when describing with reference to the attached drawings, identical components will be assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted. When describing embodiments, if a detailed description of a related known technology is judged to unnecessarily obscure the gist of the embodiment, the detailed description will be omitted.
[0044] Additionally, terms such as first, second, A, B, (a), (b), etc. may be used to describe components of the embodiments. These terms are only intended to distinguish the components from other components, and the nature, order, or sequence of the components are not limited by the terms. When a component is described as being "connected," "coupled," or "connected" to another component, it should be understood that the component may be directly connected or connected to the other component, but another component may also be "connected," "coupled," or "connected" between each component.
[0045] Components included in one embodiment and components with common functions will be described using the same names in other embodiments. Unless otherwise stated, the descriptions given in one embodiment may also apply to other embodiments, and detailed descriptions will be omitted to the extent of overlap.
[0046] Hereinafter, a translation device using gestures according to one embodiment of the present invention, its operation method, and a recording medium will be described in detail with reference to the attached drawings 1 to 21.
[0047] It should be understood that the blocks and combinations of flowcharts in each flowchart can be performed by one or more computer programs containing instructions. These one or more computer programs may be stored in their entirety on a single memory device, or different portions of the program may be stored in multiple different memory devices.
[0048] Any of the functions or operations described in the present disclosure may be processed by a single processor or a combination of multiple processors. The single processor or combination of processors is a circuit that performs processing, and may include an application processor (AP, e.g., a central processing unit (CPU)), a communication processor (CP, e.g., a modem), a graphics processing unit (GPU), a neural processing unit (NPU, e.g., an artificial intelligence (AI) chip), a Wi-Fi chip, a Bluetooth chip, a global positioning system (GPS) chip, a near field communication (NFC) chip, a connection chip, a sensor controller, a touch controller, a fingerprint sensor controller, a display driver integrated circuit (IC), an audio codec (CODEC) chip, a universal serial bus (USB) controller, a camera controller, an image processing IC, a microprocessor unit (MPU), a system on a chip (SoC), and other ICs.
[0049] FIG. 1 is a diagram illustrating the configuration of an electronic device according to an embodiment of the present disclosure.
[0050] Referring to FIG. 1, an electronic device (100) may be configured to include a processor (110), an inertial sensor (120), a microphone (130), a memory (140), a display (150), and a speaker (160).
[0051] The inertial sensor (120) includes an acceleration sensor and a gyro sensor, and can acquire sensor data including 3-axis acceleration data and 3-axis gyro data.
[0052] The microphone (130) can receive the user's voice or the other party's voice. At this time, the microphone (130) can be composed of multiple microphones.
[0053] The memory (140) can store various data used by at least one component of the electronic device (100). The data can include, for example, input data or output data for software and commands related thereto. The memory (140) can include volatile memory or non-volatile memory.
[0054] The display (150) can visually provide information to an external party (e.g., a user) of the electronic device (100). In the present disclosure, the display (150) can output translated text.
[0055] The speaker (160) can output an audio signal to the outside of the electronic device (100). The speaker (160) can be used for general purposes, such as multimedia playback or recording playback. In addition, in the present disclosure, the speaker (160) can output a converted audio signal corresponding to the translated text.
[0056] Meanwhile, in Fig. 1, the display (150) and speaker (160) may be implemented as external devices and thus omitted.
[0057] The processor (110) can control the operations of the electronic devices of FIGS. 1 to 21 by executing instructions stored in the memory (140). For example, the processor (110) can correspond to a plurality of processors that collectively perform a plurality of operations by dividing them among the processors.
[0058] The processor (110) determines the gesture of the electronic device (100) and the state (e.g., posture, direction, location) of the electronic device (100) through sensor data, and when a gesture requesting translation is detected, it determines whether the screen direction of the electronic device (100) is directed toward the user or toward the other party through the state (e.g., posture, direction, location) of the electronic device (100), and determines whether acquired voice data exists, and if acquired voice data exists, it translates the acquired voice data into the user's language or the other party's language, which is the language used by the other party, according to the screen direction of the electronic device (100), and controls the output of the translated text. At this time, the processor (110) may be composed of a plurality of processors. In addition, the acquired voice data may be the user's voice or the other party's voice. At this time, the gesture requesting translation may be a gesture that replaces the detection of the end of the user's speech, called voice activity detection (VAD) or end point detection (EPD). That is, the gesture requesting a translation may be the gesture that determines the end of the utterance.
[0059] Meanwhile, when outputting the translated text, the processor (110) measures the distance between the electronic device (100) and the user if the translated text is in the other party's language, measures the distance between the electronic device (100) and the other party if the translated text is in the user's language, and adjusts the size of the translated text in consideration of the measured distance and displays it, or adjusts the volume size of the converted audio signal corresponding to the translated text in consideration of the measured distance and outputs it.
[0060] At this time, the processor (110) can measure the distance to the user or the distance to the other party through a distance detection sensor. Alternatively, the processor (110) can measure the distance to the user or the other party through a time difference of arrival (TDoA) method using the user's voice or the other party's voice previously received through at least two microphones.
[0061] Additionally, if the translated text is not output to the display (150) at once, the processor (110) can output the translated text to the display (150) in a sliding manner.
[0062] The specific operation of the processor (110) is described in more detail below with reference to FIG. 2.
[0063] FIG. 2 is a diagram illustrating the operation of a processor of an electronic device according to an embodiment of the present disclosure.
[0064] Referring to FIG. 2, in operation 210, the processor (110) can perform preprocessing on sensor data using a filter such as a low pass filter (LPF) or a high pass filter (HPF) to remove noise from the sensor data.
[0065] In operation 220, the processor (110) calculates the amount of change in pitch using preprocessed sensor data, and calculates the magnitude of the three axes of acceleration to determine whether the electronic device (100) has stopped, thereby determining the state (e.g., posture, direction, position) of the electronic device (100). That is, the processor (110) can determine whether the electronic device (100) is facing the user or the opponent by distinguishing whether the range of the direction measured using the sensor data is included in the preset user direction or the preset opponent direction.
[0066] In operation 222, the processor (110) may perform beamforming using two or more microphones (130) to detect the speaker's speech.
[0067] In operation 224, the processor (110) can detect when the user performs a specific action on the electronic device (100) and stops, and determine whether a gesture requesting translation is detected. At this time, the specific action may be, for example, if the electronic device (100) is a smartwatch, an action of turning the wrist to change the direction of the screen of the electronic device (100) from the user's direction to the other party's direction, or an action of turning the wrist to change the direction of the screen from the other party's direction to the user's direction.
[0068] In operation 224, the processor (110) can determine that a gesture requesting translation has been detected when the pitch change amount is greater than or equal to a first threshold and the magnitude of the acceleration three-axis is less than or equal to a second threshold.
[0069] In operation 226, the processor (110) can switch the language model according to the direction of the screen.
[0070] In operation 230, the processor (110) can switch the screen so that the translated text is output in a direction that can be read by a person viewing the screen, depending on the direction of the screen.
[0071] In operation 240, the processor (110) can preprocess a voice signal. At this time, the preprocessing can improve the quality of the voice signal by performing filtering, noise removal, and frequency conversion.
[0072] In operation 250, when a gesture requesting translation is detected, the processor (110) checks whether acquired voice data exists, and if acquired voice data exists, the acquired voice data can be translated into the user's language or the other party's language depending on the screen orientation of the electronic device (100).
[0073] In operation 250, if the confirmed screen direction is toward the user, the processor (110) can translate the acquired voice of the other party into the user's language. In addition, if the confirmed screen direction is toward the other party, the processor (110) can translate the acquired voice of the user into the other party's language.
[0074] In operation 250, the processor (110) can receive the user's voice or the other party's voice after controlling the output of the translated text, or if the acquired voice data does not exist. More specifically, the processor (110) can receive the user's voice if the screen direction is toward the user, and can receive the other party's voice if the screen direction is toward the other party.
[0075] In operation 260, the processor (110) can output translated text.
[0076] In operation 260, the processor (110) may output the translated text through the display (150), or may convert the translated text into voice and output it through the speaker (160). Alternatively, the processor (110) may output the translated text through the display (150) while simultaneously converting it into voice and outputting it through the speaker (160).
[0077] In operation 260, when the processor (110) outputs the translated text through the display (150), the translated text can be output in a direction that can be read by a person viewing the display (150) through the screen transition of operation 230.
[0078]
[0079] FIG. 3 is a diagram illustrating an operation of translating voice in a processor of an electronic device according to an embodiment of the present disclosure.
[0080] Referring to FIG. 3, in operation 310, the processor (110) may perform automatic speech recognition (ASR) on the received voice signal to convert it into text. Before processing operation 310, the processor (110) may perform voice activity detection (VAD) to detect the end of the utterance. Voice activity detection is also called end point detection (EPD), and in the present disclosure, voice activity detection may be replaced by a gesture requesting translation.
[0081] In operation 320, the processor (110) performs a translation model on the converted text to determine the speaker's speech intention, translates the converted text into the user's language if the received voice signal is the other party's voice, translates the converted text into the other party's language if the received voice signal is the user's voice, and outputs the translated text result. At this time, the translation model may be composed of an LLM (Large language model) and an sLLM (small Large language model).
[0082] Although Figure 3 illustrates the 310 and 320 actions separately, the ASR and translation models can be combined into a single model. A model combining the ASR and translation models can be a S2ST (Speech-to-Speech Translation) model.
[0083] In operation 330, the processor (110) can apply TTS (text-to-speech) to the translated text result to change it into an audio signal.
[0084] Although the execution of the LLM (large language model) in this disclosure is described as an operation of the processor (110), it may also be implemented in a manner of transmitting the converted text to an externally located LLM (large language model) server (e.g., server (1808) of FIG. 18) and receiving the translated text result from the LLM (large language model) server (e.g., server (1808) of FIG. 18).
[0085] LLM can be referred to as a language model comprised of an artificial neural network pre-trained on a massive amount of text data. Compared to conventional general language models, LLM can contain more than 10 times as many parameters (e.g., more than 100 billion parameters). LLM can utilize a transformer artificial neural network structure based on an attention mechanism. The attention mechanism is a technique that helps an artificial intelligence model focus on important parts of input data. The attention mechanism can be used to predict output data by predicting the degree to which at least a portion of time-series input data (e.g., input data such as voice or video, or input data of some layers of a neural network) contributes to the intermediate or final output of the neural network. The recurrent neural network (RNN) structure, which sequentially processes each element of a sequence, has poor prediction performance when there is information dependency between long time series distances, but the attention mechanism can consider information dependency between long time series distances by controlling the degree of weight concentration within the overall (or partial) context of the input data.
[0086] For example, an LLM may include an encoder-decoder structure. The encoder may process input data and output compressed information (e.g., an attention mechanism), and the decoder may process the compressed information and output token-based data. The encoder and decoder may each include an independent attention network, and may also include a cross-attention network connecting the encoder and decoder.
[0087] For example, LLM can be trained in two stages: pre-training and fine-tuning. Pre-training involves training the LLM to process large amounts of text data and acquire general linguistic knowledge. For example, this could involve self-supervised learning, where the LLM uses previous word sequences in a text sequence to predict the next word. Fine-tuning involves training the LLM to be suitable for a specific domain (e.g., chatbot, translation, summarization, Q&A) or task. This can be achieved by further supervised learning (or adaptive learning) based on the pre-trained model using datasets tailored to the domain's purpose. LLM can perform tasks with text inputs containing natural language, called prompts. For example, LLM can include BERT (bidirectional encoder representations from transformer) and GPT (generative pre-trained transformer). The term "LLM" can refer to the neural network model itself, but can also refer to the model of an LLM-based application (e.g., chatbot, translation, summarization, text classification, sentence generation). For example, an LLM-based chatbot such as chatGPT can also be referred to as an LLM. "LLM" can also include an inference engine that utilizes the LLM neural network model. For example, "entering an input prompt into an LLM" can be referred to as "entering an input prompt into an LLM-based inference engine."
[0088] sLLM (small Large Language Model) refers to a small-scale language model, which can refer to models that use relatively small amounts of training data or are not large in size. While LLMs typically refer to models with a large number of parameters, SLLMs can encompass models with a smaller number of parameters or a simpler structure.
[0089]
[0090] Hereinafter, the method according to the present disclosure configured as above will be described with reference to the drawings below.
[0091] FIG. 4 is a flowchart illustrating an operation of performing translation according to gestures and screen orientation in an electronic device according to an embodiment of the present disclosure.
[0092] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0093] According to one embodiment, operations 410 to 480 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).
[0094] Referring to FIG. 4, according to one embodiment, in operation 410, the electronic device (100) may acquire sensor data. The acquired sensor data may be three-axis acceleration data and three-axis gyro data.
[0095] According to one embodiment, in operation 420, the electronic device (100) can determine a gesture of the electronic device (100) and a state (e.g., posture, direction, position) of the electronic device (100) through sensor data. For example, the sensor data can be used to calculate a pitch change amount, and the magnitude of the three axes of acceleration for determining whether the movement of the electronic device (100) has stopped can be calculated to determine the state of the electronic device (100).
[0096] According to one embodiment, in operation 430, the electronic device (100) can determine whether a gesture requesting translation is detected. The gesture requesting translation may be detected when the user performs a specific action with the electronic device (100) and stops, and may be detected when the amount of pitch change is greater than or equal to a first threshold and the magnitude of the acceleration triad is less than or equal to a second threshold.
[0097] According to one embodiment, if a gesture requesting translation is not detected as a result of the confirmation of operation 430, the electronic device (100) may return to operation 410 and repeat the series of operations.
[0098] According to one embodiment, if a gesture requesting translation is detected as a result of the confirmation of operation 430, in operation 440, the electronic device (100) can determine whether the screen direction of the electronic device (100) is toward the user or toward the other party through the state of the electronic device (100) (e.g., posture, direction, location).
[0099] According to one embodiment, in operation 450, the electronic device (100) can check whether acquired voice data exists. At this time, the acquired voice data may be the user's voice or the other party's voice.
[0100] According to one embodiment, if there is acquired voice data as a result of the confirmation of operation 450, in operation 460, the electronic device (100) can translate the acquired voice data into the user's language or the other party's language according to the screen orientation of the electronic device (100).
[0101] More specifically, according to one embodiment, in operation 440, if the confirmed screen direction is toward the user, in operation 460, the electronic device (100) can translate the acquired voice of the other party into the user's language. In operation 440, if the confirmed screen direction is toward the other party, in operation 460, the electronic device (100) can translate the acquired user voice into the other party's language.
[0102] According to one embodiment, in operation 470, the electronic device (100) may output the text translated in operation 460. At this time, the electronic device (100) may output the translated text on the screen, or may convert the translated text into voice and output it. Alternatively, the electronic device (100) may output the translated text on the screen while simultaneously converting it into voice and outputting it.
[0103] At this time, when the electronic device (100) outputs the translated text on the screen, it can output the translated text in a direction that can be read by the person viewing the screen. That is, when the direction of the electronic device (100) is toward the other party, the electronic device (100) can output the translated text by rotating it 180 degrees, thereby providing the other party with the translated text in a more readable manner.
[0104] According to one embodiment, after outputting the translated text in operation 470, or if the acquired voice data does not exist as a result of the confirmation in operation 450, in operation 480, the electronic device (100) can receive the user's voice or the other party's voice.
[0105] More specifically, according to one embodiment, in operation 440, if the confirmed screen direction is toward the user, in operation 480, the electronic device (100) can receive the user's voice. In operation 440, if the confirmed screen direction is toward the other party, in operation 480, the electronic device (100) can receive the other party's voice. When receiving the user's voice or the other party's voice in operation 480, the electronic device (100) can perform beamforming using two or more microphones, and can reinforce the user's voice or the other party's voice by making it louder and clearer according to the screen direction, and can minimize background noise of the reinforced voice signal. That is, if the screen direction is toward the user, the electronic device (100) can remove background noise through beamforming and make the user's voice louder and clearer to increase the clarity of the user's voice, and if the screen direction is toward the other party, the electronic device (100) can remove background noise through beamforming and make the other party's voice louder and clearer to increase the clarity of the other party's voice. At this time, when selecting a voice to be reinforced, the electronic device (100) may select the voice according to the screen direction, but may also identify a specific speech pattern or voice characteristic and detect the speaker's speech through a speaker detection algorithm.
[0106] Meanwhile, when outputting the translated text in operation 470, the electronic device (100) measures the distance from the user if the translated text is in the other party's language, measures the distance from the other party if the translated text is the user's voice, and adjusts the size of the translated text in consideration of the measured distance and displays it, or adjusts the volume size of the converted audio signal corresponding to the translated text in consideration of the measured distance and outputs it.
[0107] At this time, the electronic device (100) can measure the distance to the user or the distance to the other party through a distance detection sensor. Alternatively, the electronic device (100) can measure the distance to the user or the other party through a time difference of arrival (TDoA) method using the user's voice or the other party's voice previously received through at least two microphones.
[0108] Additionally, if the translated text is not output on the screen at once, the electronic device (100) can output the translated text by sliding it on the screen.
[0109]
[0110] FIG. 5 is a flowchart illustrating an operation of performing translation according to gestures and screen orientation in an electronic device according to an embodiment of the present disclosure.
[0111] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0112] According to one embodiment, operations 510 to 534 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).
[0113] Referring to FIG. 5, according to one embodiment, in operation 510, the electronic device (100) may acquire sensor data. The acquired sensor data may be three-axis acceleration data and three-axis gyro data.
[0114] According to one embodiment, in operation 512, the electronic device (100) can determine a gesture of the electronic device (100) and a state (e.g., posture, direction, location) of the electronic device (100) through sensor data. For example, the sensor data can be used to calculate a pitch change amount, and the magnitude of three axes of acceleration for determining whether the electronic device (100) has stopped can be calculated to determine the state (e.g., posture, direction, location).
[0115] According to one embodiment, in operation 514, the electronic device (100) can determine whether a gesture requesting translation is detected. The gesture requesting translation may be detected when the user performs a specific action with the electronic device (100) and stops, and may be detected when the amount of pitch change is greater than or equal to a first threshold and the magnitude of the acceleration triad is less than or equal to a second threshold.
[0116] According to one embodiment, if a gesture requesting translation is not detected as a result of the confirmation of operation 514, the electronic device (100) may return to operation 510 and repeat the series of operations.
[0117] According to one embodiment, if a gesture requesting translation is detected as a result of the confirmation of operation 514, in operation 516, the electronic device (100) can check whether the screen direction of the electronic device (100) is in the direction of the user through the state of the electronic device (100) (e.g., posture, direction, position).
[0118] According to one embodiment, if the screen direction of the electronic device (100) is oriented toward the user as a result of the confirmation in operation 516, in operation 518, the electronic device (100) can check whether acquired voice data exists. At this time, the acquired voice data may be the voice of the other party.
[0119] According to one embodiment, if there is acquired voice data as a result of the confirmation of operation 518, in operation 520, the electronic device (100) can translate the acquired voice data into the user's language.
[0120] And, in operation 522, the electronic device (100) can output the text translated in operation 520. At this time, the electronic device (100) can output the translated text on the screen, or can convert the translated text into voice and output it. Alternatively, the electronic device (100) can output the translated text on the screen while simultaneously converting it into voice and outputting it.
[0121] According to one embodiment, after outputting the translated text in operation 522, or if the acquired voice data does not exist as a result of the confirmation in operation 518, the electronic device (100) may receive the user's voice in operation 524. Then, the electronic device (100) may perform operation 524 and return to operation 510 to repeat the series of processes.
[0122] According to one embodiment, if the screen direction of the electronic device (100) is not in the direction of the user as a result of the confirmation of operation 516, it can be confirmed in operation 526 whether the screen direction of the electronic device (100) is in the direction of the other party.
[0123] According to one embodiment, if the screen direction of the electronic device (100) is not the direction of the other party as a result of the confirmation of operation 526, the electronic device (100) can return to operation 510 and repeat a series of operations.
[0124] According to one embodiment, if the screen direction of the electronic device (100) is oriented toward the other party as a result of the confirmation of operation 526, the electronic device (100) can confirm whether acquired voice data exists. In this case, the acquired voice data may be the user's voice.
[0125] According to one embodiment, if the acquired voice data exists as a result of the confirmation of operation 528, in operation 530, the electronic device (100) can translate the acquired voice data into the other party's language.
[0126] And, according to one embodiment, in operation 532, the electronic device (100) may output the text translated in operation 530. At this time, the electronic device (100) may output the translated text on the screen, or may convert the translated text into voice and output it. Alternatively, the electronic device (100) may output the translated text on the screen while simultaneously converting it into voice and outputting it.
[0127] At this time, when the electronic device (100) outputs the translated text on the screen, it can provide the translated text to the other party by rotating the translated text 180 degrees so that the other party can read it more easily.
[0128] According to one embodiment, after outputting the translated text in operation 532, or if the acquired voice data does not exist as a result of the confirmation in operation 528, the electronic device (100) can receive the other party's voice in operation 534. Then, the electronic device (100) can perform operation 534 and return to operation 510 to repeat the series of processes.
[0129]
[0130] FIG. 6 is a flowchart illustrating an operation of checking a screen orientation in an electronic device according to an embodiment of the present disclosure.
[0131] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0132] According to one embodiment, operations 610 to 670 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).
[0133] Referring to FIG. 6, according to one embodiment, in operation 610, the electronic device (100) may acquire sensor data. At this time, the sensor data may be 3-axis acceleration data of an acceleration sensor and 3-axis gyro data of a gyro sensor.
[0134] According to one embodiment, in operation 620, the electronic device (100) can remove noise from sensor data. The electronic device (100) can remove noise from sensor data using a filter such as a low pass filter (LPF) or a high pass filter (HPF).
[0135] According to one embodiment, in operation 630, the electronic device (100) can calculate the pitch change amount and the magnitude of the acceleration three axes. At this time, the pitch change amount means the rotation amount of the electronic device (100) based on the axis connecting the 3 o'clock and 9 o'clock directions of the electronic device (100), and can be determined through the differential sum or dispersion for a certain period of time.
[0136] According to one embodiment, in operation 640, the electronic device (100) can determine whether the pitch change amount is greater than or equal to a first threshold value.
[0137] According to one embodiment, if the pitch change amount as a result of the verification of operation 640 is greater than or equal to the first threshold value, in operation 650, the electronic device (100) can verify whether the magnitude of the acceleration three-axis is less than or equal to the second threshold value.
[0138] According to one embodiment, if the pitch change amount is not greater than the first threshold as a result of the confirmation of operation 640 or the magnitude of the acceleration triaxiality is not less than the second threshold as a result of the confirmation of operation 650, the electronic device (100) determines that a gesture requesting translation has not been detected, and returns to operation 610 to repeat a series of operations.
[0139] According to one embodiment, if the magnitude of the acceleration triaxiality is less than or equal to the second threshold as a result of the verification of operation 650, then in operation 660, the electronic device (100) can determine the state (e.g., posture, direction, position) of the electronic device (100). That is, when there is no movement of the electronic device (100), the electronic device (100) can determine the final pitch angle of the electronic device (100).
[0140] According to one embodiment, in operation 670, the electronic device (100) can determine whether the screen direction of the electronic device (100) is toward the user or toward the other party through the status determination result of operation 650.
[0141] An example of determining whether the direction is user or opponent's direction using sensor data is described below with reference to Figure 9.
[0142] FIG. 9 is a diagram illustrating an example of determining a user direction and an opponent direction based on a sensor measurement value in an electronic device according to an embodiment of the present disclosure.
[0143] Referring to FIG. 9, 910 is a graph showing changes in 3-axis acceleration data as the screen orientation changes when the electronic device (100) is a smart watch. In the graph of 910, the x-axis is a sample value, and the y-axis is am.
[0144] 920 is a graph showing changes in 3-axis gyro data as the screen orientation changes when the electronic device (100) is a smartwatch. In the graph of 920, the x-axis is a sample value and the y-axis is deg / s.
[0145] When examining the section (930) where the state (e.g., posture, direction, position) is switched and the direction of the screen is changed, it can be confirmed that the change in the 3-axis acceleration data and the change in the 3-axis gyro data are greatly changed.
[0146] And, when the user direction (942, 946) and the opponent direction (944) are maintained, it can be confirmed that the change in the 3-axis acceleration data and the change in the 3-axis gyro data are not large.
[0147]
[0148] FIG. 7 is a flowchart illustrating an operation of detecting a speaker's speech in an electronic device according to an embodiment of the present disclosure.
[0149] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0150] According to one embodiment, operations 710 to 750 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).
[0151] Referring to FIG. 7, according to one embodiment, in operation 710, the electronic device (100) may acquire audio signals from two or more microphones. At this time, the acquired audio signals may be multi-channel.
[0152] According to one embodiment, in operation 720, the electronic device (100) may preprocess the acquired audio signal. At this time, the preprocessing may improve the quality of the audio signal by performing filtering, noise removal, and frequency conversion.
[0153] According to one embodiment, in operation 730, the electronic device (100) may emphasize an audio signal in a primary sound direction through a beamforming algorithm. In this case, emphasizing an audio signal in a primary sound direction may mean removing background noise and increasing the size and clarity of the audio signal in the primary sound direction.
[0154] In one embodiment, in operation 740, the electronic device (100) can remove background noise based on the emphasized audio signal. That is, the electronic device (100) can emphasize speech in the main sound direction and minimize background noise.
[0155] According to one embodiment, in operation 750, the electronic device (100) can detect the speaker's speech from the processed audio signal. More specifically, the electronic device (100) can identify a specific speech pattern or voice characteristic from the processed audio signal through a speaker detection algorithm and detect the speaker's speech.
[0156]
[0157] FIG. 8 is a flowchart illustrating an operation of translating a voice signal in an electronic device according to an embodiment of the present disclosure.
[0158] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0159] According to one embodiment, operations 810 to 860 may be understood to be performed in a processor (e.g., processor (110) of FIG. 1) of an electronic device (e.g., electronic device (100) of FIG. 1).
[0160] Referring to FIG. 8, according to one embodiment, in operation 810, the electronic device (100) may acquire a voice signal. At this time, the electronic device (100) may determine how much of the input audio to utilize as voice data using a user's gesture requesting translation. For example, the electronic device (100) may determine that a gesture requesting translation has been detected by detecting that the screen direction of the electronic device (100) changes from the user's direction to the other party's direction or from the other party's direction to the user's direction. After detecting a gesture requesting translation, the electronic device (100) may not utilize the voice input through the microphone in subsequent processing operations 820 to 860 before outputting the translated text.
[0161] According to one embodiment, in operation 820, the electronic device (100) may perform preprocessing of the collected voice signal. At this time, the electronic device (100) may improve the quality of the voice signal by preprocessing the voice signal through at least one of noise removal, filtering, and normalization.
[0162] According to one embodiment, in operation 830, the electronic device (100) can extract voice frequency features. At this time, the electronic device (100) can extract the frequency characteristics of the voice using a frequency transformation technique such as Fourier transform or MFCC (mel-frequency cepstral coefficients).
[0163] According to one embodiment, in operation 840, the electronic device (100) may convert speech into text by applying a language model to the frequency characteristics of the extracted speech. At this time, the electronic device (100) may select and apply a language model corresponding to the speech, taking into account the screen orientation.
[0164] According to one embodiment, in operation 850, the electronic device (100) may translate text into a target language, which is a language to be translated, using a translation model. At this time, the electronic device (100) may translate the input text into the target language based on language patterns learned through machine translation and the translation model. In FIG. 8, operations 840 and 850 are described separately, but the language model and the translation model may be combined into one. A model in which the language model and the translation model are combined into one may be S2ST (Speech-to-Speech Translation).
[0165] According to one embodiment, in operation 860, the electronic device (100) may output the interpreted result. At this time, the electronic device (100) may output the interpreted result in text or voice through a screen or speaker.
[0166]
[0167] Then, an example of an electronic device for translating using the gesture of the present disclosure being applied to a smartwatch is described below through FIGS. 10 to 14.
[0168] FIG. 10 is a drawing illustrating an example of providing a user's voice translated into the other party's language according to an embodiment of the present disclosure.
[0169] Referring to FIG. 10, when the electronic device (100) is a smartwatch, the electronic device (100) receives the user's voice "Would you like a cup of coffee now?" (1012) from the user's direction (1010), and when the screen direction of the electronic device (100) is changed to the other party's direction (1020) through a wrist-twisting gesture, the electronic device can output "Would you like a cup of coffee now?" (1022) by translating the received user's voice into the other party's language.
[0170] In FIG. 10, the electronic device (100) may output the user's voice as text through the screen of the electronic device (100) when the screen direction of the electronic device (100) is toward the user (1010), thereby allowing the user to confirm that the user's voice is being input properly.
[0171]
[0172] FIG. 11 is a drawing illustrating an example of providing a translation of the other party's voice into the user's language according to an embodiment of the present disclosure.
[0173] Referring to FIG. 11, when the electronic device (100) is a smartwatch, the electronic device (100) receives the other party's voice "Sure" (1112) from the other party's direction (1110), and when the screen direction of the electronic device (100) is changed to the user's direction (1120) through a wrist-twisting gesture, the electronic device can output "Like" (1122), which is a translation of the received other party's voice into the user's language.
[0174] In FIG. 11, when the screen direction of the electronic device (100) is toward the other party (1110), the electronic device (100) can output the voice of the other party as text through the screen of the electronic device (100), thereby allowing the other party to confirm that the other party's voice is being input properly.
[0175]
[0176] FIG. 12 is a drawing illustrating an example of outputting in a direction that can be read by a person viewing the screen according to the screen orientation according to an embodiment of the present disclosure.
[0177] Referring to FIG. 12, when the electronic device (100) is a smartwatch, when the screen direction is changed to the user direction (1210) in which the screen direction is toward the user, the electronic device (100) can output the translated version of the other party's voice, "Hello! The delivery has arrived. Could you please sign it?" (1212), as "hello! The delivery has arrived. Could you please sign it?" (1214), in a direction that is easy for the user to read through the screen of the electronic device (100).
[0178] In addition, when the screen direction of the electronic device (100) is changed to the direction of the other party (1220), which is the direction facing the other party, the electronic device (100) can output the user's voice, "Yes, Thank you." (1222), translated as "Yes, thank you." (1224), through the screen of the electronic device (100) in a direction that is easy for the other party to read (for example, rotated by about 180 degrees).
[0179]
[0180] FIG. 13 is a drawing illustrating an example of adjusting the font size according to the distance between an electronic device and a counterpart according to an embodiment of the present disclosure.
[0181] Referring to FIG. 13, the electronic device (100) can adaptively adjust the size of translated text displayed on the screen of the electronic device (100) and the sound of the speaker according to the distance between the electronic device (100) and the other party.
[0182] When the distance between the other party and the electronic device (100) is close (e.g., about 30 cm) (1310), the electronic device (100) can output letters in a size usually provided (a preset size) (e.g., font size 7), and in the case of audio, can provide it at speaker level 3 or so.
[0183] However, if the distance between the counterpart and the electronic device (100) is far (e.g., about 100 cm) (1320), the electronic device (100) may output letters in a preset size (e.g., font size 12) according to the distance, and provide audio at speaker level 7 or so. In addition, the electronic device (100) may also make the font bold or change the color of the font to improve visibility.
[0184] That is, when the distance between the other party and the electronic device (100) is far (1320) than when the distance between the other party and the electronic device (100) is close (1310), the translated content can be provided to the other party with a larger font size and louder audio.
[0185] If the electronic device (100) cannot output the translated content on one screen due to the size of the translated text being adjusted to a larger size or if the user wishes to output the summarized content (1330), the electronic device (100) may summarize the translated content and output the summarized content. For example, the electronic device (100) may summarize the translated content (Would you like a cup of coffee now?) or extract keywords and output the summarized content (coffee?) or keywords (coffee?).
[0186]
[0187] FIG. 14 is a drawing illustrating an example of sliding and outputting letters according to the distance between an electronic device and a counterpart according to an embodiment of the present disclosure.
[0188] Referring to FIG. 14, if the distance between the counterpart and the electronic device (100) is far and the text size is set to be large, and the translated content (Would you like a cup of coffee now?) is not displayed all at once on the screen of the electronic device (100), the translated content can be displayed in a sliding manner, as shown in operation 1410 to operation 1420.
[0189]
[0190] FIG. 15 is a diagram illustrating an example of receiving a user's voice when the other party is present, according to an embodiment of the present disclosure.
[0191] Referring to FIG. 15, when the electronic device (100) is a smartwatch, if the electronic device (100) receives the voice of the user (1510) “Would you like a cup of coffee now?” (1530) from a direction facing the user (1510), the electronic device (100) can output the text “Would you like a cup of coffee now?” (1540) corresponding to the voice received through the screen of the electronic device (100). At this time, the electronic device (100) can output “Would you like a cup of coffee now?” (1540) in a direction that is easy for the user (1510) to read.
[0192]
[0193] FIG. 16 is a diagram illustrating an example of providing a translation of a user's voice into the language of a counterpart when the counterpart is present, according to an embodiment of the present disclosure.
[0194] When the electronic device (100) is a smartwatch, when the screen direction of the electronic device (100) changes from a direction facing the user (1510) to a direction facing the other party (1520) (1020) through a gesture of extending the wrist, the electronic device (100) can output the received user's voice, "Would you like a cup of coffee now?" (1530), translated into the other party's (1520) language, "Would you like a cup of coffee now?" (1640). At this time, the electronic device (100) can output "Would you like a cup of coffee now?" (1640) in a direction that is easy for the other party (1520) to read.
[0195] Meanwhile, the electronic device (100) can detect a wrist-extending gesture through a change in the acceleration sensor included in the inertial sensor (120).
[0196] In the case of FIGS. 15 and 16, the user (1510) and the other party (1520) are not standing facing each other but are standing sideways, so the directions in which the user (1510) and the other party (1520) can read comfortably may be the same.
[0197]
[0198] Meanwhile, the electronic device (100) of FIG. 1 may be configured in the form of an electronic device (1700) as in FIG. 17 below, may be configured in the form of an electronic device (1801) in a network environment as in FIG. 18, or may be configured in the form of a smart watch as in FIG. 19 to FIG. 21.
[0199] FIG. 17 is a block diagram schematically illustrating the structure of an electronic device according to an embodiment of the present disclosure.
[0200] Referring to FIG. 17, an electronic device (1700) may be configured to include a processor (1710) and a memory (1720).
[0201] The memory (1720) can store various data used by at least one component (e.g., the processor (1710)) of the electronic device (1700). The data can include, for example, input data or output data for software and commands related thereto. The memory (1720) can include volatile memory or non-volatile memory. In this case, the memory (1720) can have a configuration corresponding to the memory (140) of FIG. 1.
[0202] The processor (1710) can control the overall operation of the electronic device (1700). When a gesture requesting translation is detected, the processor (1710) determines whether the screen direction of the electronic device (1700) is toward the user or the other party, determines whether acquired voice data exists, and if acquired voice data exists, translates the acquired voice data into the user's language or the other party's language according to the screen direction of the electronic device (1700), and controls the output of the translated text. At this time, the processor (1710) may be composed of a plurality of processors. The acquired voice data may be the user's voice or the other party's voice. At this time, the processor (1710) may have a configuration corresponding to the processor (110) of FIG. 1.
[0203]
[0204] FIG. 18 is a block diagram illustrating an electronic device within a network environment according to an embodiment of the present disclosure.
[0205] Referring to FIG. 18, in a network environment (1800), an electronic device (1801) may communicate with an electronic device (1802) via a first network (1898) (e.g., a short-range wireless communication network), or may communicate with an electronic device (1804) or a server (1808) via a second network (1899) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1801) may communicate with the electronic device (1804) via the server (1808). According to one embodiment, the electronic device (1801) may include a processor (1820), a memory (1830), an input module (1850), an audio output module (1855), a display module (1860), an audio module (1870), a sensor module (1876), an interface (1877), a connection terminal (1878), a haptic module (1879), a camera module (1880), a power management module (1888), a battery (1889), a communication module (1890), a subscriber identification module (1896), or an antenna module (1897). In some embodiments, the electronic device (1801) may omit at least one of these components (e.g., the connection terminal (1878)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1876), camera module (1880), or antenna module (1897)) may be integrated into a single component (e.g., display module (1860)).
[0206] The processor (1820) may, for example, execute software (e.g., a program (1840)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1801) connected to the processor (1820) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operation, the processor (1820) may store commands or data received from other components (e.g., a sensor module (1876) or a communication module (1890)) in a volatile memory (1832), process the commands or data stored in the volatile memory (1832), and store the resulting data in a non-volatile memory (1834). At this time, the processor (1820) may have a configuration corresponding to the processor (110) of FIG. 1.
[0207] According to one embodiment, the processor (1820) may include a main processor (1821) (e.g., a central processing unit or an application processor) or an auxiliary processor (1823) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1821). For example, when the electronic device (1801) includes the main processor (1821) and the auxiliary processor (1823), the auxiliary processor (1823) may be configured to use less power than the main processor (1821) or to be specialized for a given function. The auxiliary processor (1823) may be implemented separately from the main processor (1821) or as a part thereof.
[0208] The auxiliary processor (1823) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1860), the sensor module (1876), or the communication module (1890)) of the electronic device (1801), for example, on behalf of the main processor (1821) while the main processor (1821) is in an inactive (e.g., sleep) state, or together with the main processor (1821) while the main processor (1821) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1823) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1880) or a communication module (1890)). In one embodiment, the auxiliary processor (1823) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1801) where the artificial intelligence is performed, or can be performed through a separate server (e.g., server (1808)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0209] The memory (1830) can store various data used by at least one component (e.g., the processor (1820) or the sensor module (1876)) of the electronic device (1801). The data can include, for example, software (e.g., the program (1840)) and input data or output data for commands related thereto. The memory (1830) can include a volatile memory (1832) or a non-volatile memory (1834). In this case, the memory (1830) can have a configuration corresponding to the memory (140) of FIG. 1.
[0210] The program (1840) may be stored as software in memory (1830) and may include, for example, an operating system (1842), middleware (1844), or an application (1846).
[0211] The input module (1850) can receive commands or data to be used in a component (e.g., processor (1820)) of the electronic device (1801) from an external source (e.g., a user) of the electronic device (1801). The input module (1850) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen). In this case, the input module (1850) can be configured to include the microphone (130) of FIG. 1.
[0212] The audio output module (1855) can output audio signals to the outside of the electronic device (1801). The audio output module (1855) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. In one embodiment, the receiver may be implemented separately from the speaker or as part of the speaker. In this case, the audio output module (1855) may be configured to include the speaker (160) of FIG. 1.
[0213] The display module (1860) can visually provide information to an external party (e.g., a user) of the electronic device (1801). The display module (1860) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device. According to one embodiment, the display module (1860) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch. In this case, the display module (1860) may have a configuration corresponding to the display (150) of FIG. 1.
[0214] The audio module (1870) can convert sound into an electrical signal, or vice versa. According to one embodiment, the audio module (1870) can acquire sound through the input module (1850), output sound through the sound output module (1855), or an external electronic device (e.g., electronic device (1802)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1801).
[0215] The sensor module (1876) can detect the operating status (e.g., power or temperature) of the electronic device (1801) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1876) may include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor. In this case, the sensor module (1876) may be configured to include the inertial sensor (120) of FIG. 1.
[0216] The interface (1877) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1801) with an external electronic device (e.g., the electronic device (1802)). In one embodiment, the interface (1877) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0217] The connection terminal (1878) may include a connector through which the electronic device (1801) may be physically connected to an external electronic device (e.g., the electronic device (1802)). In one embodiment, the connection terminal (1878) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0218] The haptic module (1879) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1879) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0219] The camera module (1880) can capture still images and moving images. In one embodiment, the camera module (1880) may include one or more lenses, image sensors, image signal processors, or flashes.
[0220] The power management module (1888) can manage the power supplied to the electronic device (1801). According to one embodiment, the power management module (1888) can be implemented as at least a part of, for example, a power management integrated circuit (PMIC).
[0221] A battery (1889) may power at least one component of the electronic device (1801). In one embodiment, the battery (1889) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0222] The communication module (1890) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1801) and an external electronic device (e.g., electronic device (1802), electronic device (1804), or server (1808)), and the performance of communication through the established communication channel. The communication module (1890) may operate independently from the processor (1820) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1890) may include a wireless communication module (1892) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1894) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1804) via a first network (1898) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1899) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1892) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1896) to identify or authenticate the electronic device (1801) within a communication network such as the first network (1898) or the second network (1899).
[0223] The wireless communication module (1892) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1892) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1892) may support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1892) may support various requirements specified in the electronic device (1801), an external electronic device (e.g., the electronic device (1804)), or a network system (e.g., the second network (1899)). According to one embodiment, the wireless communication module (1892) can support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC implementation.
[0224] The antenna module (1897) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1897) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1897) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1898) or the second network (1899), may be selected from the plurality of antennas by, for example, the communication module (1890). A signal or power may be transmitted or received between the communication module (1890) and the external electronic device via the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1897).
[0225] According to various embodiments, the antenna module (1897) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0226] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0227] According to one embodiment, commands or data may be transmitted or received between the electronic device (1801) and an external electronic device (1804) via a server (1808) connected to a second network (1899). Each of the external electronic devices (1802 or 1804) may be the same or a different type of device as the electronic device (1801). According to one embodiment, all or part of the operations executed in the electronic device (1801) may be executed in one or more of the external electronic devices (1802 or 1804) or the server (1808). For example, when the electronic device (1801) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1801) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1801). The electronic device (1801) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1801) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1804) may include an Internet of Things (IoT) device. The server (1808) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1804) or server (1808) may be included within the second network (1899). The electronic device (1801) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0228]
[0229] FIG. 19 is a front perspective view of an electronic device according to an embodiment of the present disclosure.
[0230] FIG. 20 is a rear perspective view of an electronic device according to an embodiment of the present disclosure.
[0231] Referring to FIGS. 19 and 20 , an electronic device (1900) according to one embodiment (e.g., the electronic device (1801) of FIG. 18 ) may include a housing (1910) including a first side (or front side) (1910A), a second side (or back side) (1910B), and a side surface (1910C) enclosing a space between the first side (1910A) and the second side (1910B), and a fastening member (1950, 1960) connected to at least a portion of the housing (1910) and configured to detachably fasten the electronic device (1900) to a body part (e.g., a wrist, an ankle, etc.) of a user. In another embodiment (not shown), the housing may also refer to a structure forming a portion of the first side (1910A), the second side (1910B), and the side surface (1910C) of FIG. 19 . In one embodiment, the first side (1910A) may be formed by a front plate (1901) that is at least partially substantially transparent (e.g., a glass plate or a polymer plate comprising various coating layers). The second side (1910B) may be formed by a substantially opaque back plate (1907). The back plate (1907) may be formed of, for example, coated or colored glass, ceramic, polymer, metal (e.g., aluminum, stainless steel (STS), or magnesium), or a combination of at least two of the foregoing materials. The side surface (1910C) may be formed by a side bezel structure (or “side member”) (1906) that is coupled to the front plate (1901) and the back plate (1907) and comprises a metal and / or a polymer. In some embodiments, the back plate (1907) and the side bezel structure (1906) may be formed integrally and comprise the same material (e.g., a metal material such as aluminum). The above-mentioned bonding member (1950, 1960) can be formed of various materials and shapes.The integral and multiple unit links can be formed to be movable with each other by woven materials, leather, rubber, urethane, metal, ceramic, or a combination of at least two of the above materials.
[0232] According to one embodiment, the electronic device (1900) may include at least one of a display (1920, see FIG. 21), a microphone hole (1905) and a speaker (1908) of an audio module (1870), a sensor module (1911), a key input device (1902, 1903, 1904), and a connector hole (1909). In some embodiments, the electronic device (1900) may omit at least one of the components (e.g., the key input device (1902, 1903, 1904), the connector hole (1909), or the sensor module (1911)), or may additionally include other components.
[0233] The display (1920) may be exposed, for example, through a significant portion of the front plate (1901). The shape of the display (1920) may correspond to the shape of the front plate (1901), and may have various shapes such as a circle, an oval, or a polygon. The display (1920) may be coupled to or disposed adjacent to a touch sensing circuit, a pressure sensor capable of measuring the intensity (pressure) of a touch, and / or a fingerprint sensor.
[0234] The audio module (1870) may include a microphone hole (1905) and a speaker hole (1908). The microphone hole (1905) may have a microphone positioned therein for acquiring external sounds, and in some embodiments, multiple microphones may be positioned therein to detect the direction of sounds. The speaker hole (1908) may be used as an external speaker and a receiver for calls. In some embodiments, the speaker hole (1908) and the microphone hole (1905) may be implemented as a single hole, or a speaker may be included without the speaker hole (1908) (e.g., a piezo speaker).
[0235] The sensor module (1911) can generate an electrical signal or data value corresponding to an internal operating state of the electronic device (1900) or an external environmental state. The sensor module (1911) can include, for example, a biometric sensor module (e.g., a heart rate monitor (HRM) sensor) arranged on the second surface (1910B) of the housing (1910). The electronic device (1900) can further include at least one of a non-illustrated sensor module, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a color sensor, an infrared (IR) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0236] The sensor module (1911) may include electrode regions (1913, 1914) forming a portion of a surface of the electronic device (1900) and a biosignal detection circuit (not shown) electrically connected to the electrode regions (1913, 1914). For example, the electrode regions (1913, 1914) may include a first electrode region (1913) and a second electrode region (1914) disposed on a second surface (1910B) of the housing (1910). The sensor module (1911) may be configured such that the electrode regions (1913, 1914) obtain electrical signals from a portion of the user's body, and the biosignal detection circuit detects bioinformation of the user based on the electrical signals.
[0237] The key input devices (1902, 1903, 1904) may include a wheel key (1902) disposed on a first side (1910A) of the housing (1910) and rotatable in at least one direction, and / or a side key button (1903, 1904) disposed on a side (1910C) of the housing (1910). The wheel key may have a shape corresponding to the shape of the front plate (1901). In other embodiments, the electronic device (1900) may not include some or all of the above-mentioned key input devices (1902, 1903, 1904), and the key input devices (1902, 1903, 1904) that are not included may be implemented in another form, such as a soft key, on the display (1920). The connector hole (1909) may include another connector hole (not shown) that may accommodate a connector (e.g., a USB connector) for transmitting and receiving power and / or data with an external electronic device, and may accommodate a connector for transmitting and receiving audio signals with the external electronic device. The electronic device (1900) may further include, for example, a connector cover (not shown) that covers at least a portion of the connector hole (1909) and blocks foreign substances from entering the connector hole.
[0238] The fastening member (1950, 1960) can be removably fastened to at least a portion of the housing (1910) using a locking member (1951, 1961). The fastening member (1950, 1960) can include one or more of a fixing member (1952), a fixing member fastening hole (1953), a band guide member (1954), and a band fastening ring (1955).
[0239] The fixing member (1952) can be configured to fix the housing (1910) and the fastening members (1950, 1960) to a part of the user's body (e.g., wrist, ankle). The fastening member fastening hole (1953) can fix the housing (1910) and the fastening members (1950, 1960) to a part of the user's body in response to the fastening member (1952). The band guide member (1954) can be configured to limit the range of motion of the fastening member (1952) when the fastening member (1952) is fastened to the fastening member fastening hole (1953), thereby allowing the fastening members (1950, 1960) to be fastened in close contact with a part of the user's body. The band fastening ring (1955) can limit the range of motion of the fastening members (1950, 1960) when the fastening member (1952) and the fastening member fastening hole (1953) are fastened.
[0240]
[0241] FIG. 21 is an exploded perspective view of an electronic device according to an embodiment of the present disclosure.
[0242] Referring to FIG. 21, an electronic device (2100) (e.g., the electronic device (1801) of FIG. 18 or the electronic device (1900) of FIGS. 19 and 20) may include a side bezel structure (2110), a wheel key (2120), a front plate (1901), a display (1920), a first antenna (2150), a second antenna (2155), a support member (2160) (e.g., a bracket), a battery (2170), a printed circuit board (2180), a sealing member (2190), a rear plate (2193) (e.g., the rear plate (1907) of FIG. 19), and a fastening member (2195, 2197) (e.g., the fastening member (1950, 1960) of FIGS. 19 to 20). At least one of the components of the electronic device (2100) may be identical or similar to at least one of the components of the electronic device (1900) of FIG. 18, or FIG. 19, or FIG. 20, and any overlapping descriptions will be omitted below. The support member (2160) may be disposed inside the electronic device (2100) and connected to the side bezel structure (2110), or may be formed integrally with the side bezel structure (2110). The support member (2160) may be formed of, for example, a metal material and / or a non-metallic (e.g., a polymer) material. The support member (2160) may have a display (1920) coupled to one surface and a printed circuit board (2180) coupled to the other surface. A processor, a memory, and / or an interface may be mounted on the printed circuit board (2180). The processor may include, for example, one or more of a central processing unit, an application processor, a graphics processing unit (GPU), a sensor processor, or a communication processor.
[0243] The memory may include, for example, volatile memory or non-volatile memory. The interface may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, a secure digital (SD) card interface, and / or an audio interface. The interface may electrically or physically connect the electronic device (2100) to an external electronic device, for example, and may include a USB connector, an SD card / MMC (multimedia card) connector, or an audio connector.
[0244] The battery (2170) is a device for supplying power to at least one component of the electronic device (2100), and may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell. At least a portion of the battery (2170) may be disposed substantially on the same plane as, for example, the printed circuit board (2180). The battery (2170) may be disposed integrally within the electronic device (1900), or may be disposed detachably from the electronic device (1900).
[0245] The first antenna (2150) may be positioned between the display (1920) and the support member (2160). The first antenna (2150) may include, for example, a near field communication (NFC) antenna, a wireless charging antenna, and / or a magnetic secure transmission (MST) antenna. The first antenna (2150) may, for example, perform short-range communication with an external device, wirelessly transmit and receive power required for charging, and transmit a magnetic-based signal including a short-range communication signal or payment data. In another embodiment, the antenna structure may be formed by a portion or a combination of the side bezel structure (2110) and / or the support member (2160).
[0246] The second antenna (2155) may be positioned between the printed circuit board (2180) and the back plate (2193). The second antenna (2155) may include, for example, a near field communication (NFC) antenna, a wireless charging antenna, and / or a magnetic secure transmission (MST) antenna. The second antenna (2155) may, for example, perform short-range communication with an external device, wirelessly transmit and receive power required for charging, and transmit a magnetic-based signal including a short-range communication signal or payment data. In another embodiment, the antenna structure may be formed by a portion or a combination of the side bezel structure (2110) and / or the back plate (2193).
[0247] A sealing member (2190) may be positioned between the side bezel structure (2110) and the rear plate (2193). The sealing member (2190) may be configured to block moisture and foreign substances from entering the space surrounded by the side bezel structure (2110) and the rear plate (2193) from the outside.
[0248]
[0249] FIG. 22 is a diagram illustrating an example of a combination that provides a translation service through linkage between a wearable device that detects gestures and a mobile device according to an embodiment of the present disclosure.
[0250] Referring to FIG. 22, detection of a gesture requesting translation can be performed by various wearable devices (2211, 2212, 2213, 2214). For example, in the case of a smartwatch (2211) (e.g., Galaxy Watch), a smart ring (2212) (e.g., Galaxy Ring), or a smart band (2214) (e.g., Galaxy Fit) worn on the wrist or finger, a gesture of turning the device, a gesture of extending the hand toward the other party, or a double-tap gesture (the double-tap gesture is explained with reference to FIG. 24 below) on the hand wearing the wearable device (2211, 2212, 2213, 2214) can be set as a gesture requesting translation. In the case of a wearable device, such as a wireless earphone (2213) (e.g., Galaxy Buds), a gesture of putting it on and nodding the head or a gesture of turning it left and right can be set as a gesture requesting translation. In the case of the smart ring (2212), a gesture of tapping the ring worn on the index finger, a gesture of swiping up or down, or a gesture of snapping fingers can be set as a gesture for requesting translation. The gestures of various wearable devices (2211, 2212, 2213, 2214) can be detected through inertial sensors within the various wearable devices (2211, 2212, 2213, 2214). In this case, the gesture for requesting translation may be a gesture that replaces voice activity detection, which confirms the end of utterance.
[0251] In addition, various wearable devices (2211, 2212, 2213, 2214) can use multiple gestures to distinguish actions. For example, when displaying translation results on a linked device, a wrist-turning gesture of a smartwatch (2211) can be used as a gesture for voice activation detection and switching language models, and a double-tap gesture of the smartwatch (2211) can be used as a gesture for voice activation detection. In this way, when a user of the smartwatch (2211) speaks for a long time, if the user makes a double-tap gesture during the speech, the user's speech input up to that point can be translated and displayed through the linked mobile device (2221, 2222, 2223) to show it to the other party, and if a gesture of turning the smartwatch (2211) is made, the user's speech up to that point can be translated and the result can be output, and the translation model can be changed to interpret the other party's speech. In the case of the smart ring (2212), a gesture of swiping up or down on the smart ring (2212) worn on the index finger can perform voice activation detection and language model switching, and a gesture of tapping the smart ring (2212) can only perform voice activation detection.
[0252] A wearable device (2211, 2213, 2214) with a microphone can receive speech input from a user and the other party, interpret it, and transmit the result to a linked mobile device (2221, 2222, 2223) for output. In addition, a wearable device (2211, 2213, 2214) with a microphone can receive voice input, and translation can be performed on a linked mobile device (2221, 2222, 2223). For example, when the watch (2211) receives a voice input from a user or the other party and performs a gesture, it transmits text results or audio data and information about a language model through ASR to a linked device (2221, 2222, 2223), and the linked mobile device (2221, 2222, 2223) can interpret the received text results or audio data using an appropriate language model and output the results on the linked mobile device (2221, 2222, 2223).
[0253] In the case of a wearable device (2212) without a microphone or in a case where it is difficult to input the other party's speech (e.g., wireless earphones (2213)), the user's speech or the other party's speech can be input through a linked mobile device (2221, 2222, 2223), and the wearable device (2212) without a microphone can be used for the purpose of voice activity detection.
[0254] In the case of wireless earphones (2213), the user and the other party can wear them separately. In this case, the right wireless earphone and the left wireless earphone each receive the voice input of the user wearing them or the voice input of the other party wearing them, detect a gesture, and transmit the text result or audio data through ASR to the linked device (2221, 2222, 2223), and the linked mobile device (2221, 2222, 2223) translates the received text result or audio data using an appropriate language model, outputs the translated result through the screen of the mobile device (2221, 2222, 2223), and outputs the translated result as audio through the right wireless earphone or the left wireless earphone worn by the user or the other party.
[0255]
[0256] FIG. 23 is a diagram illustrating an example of outputting a translated result according to an embodiment of the present disclosure.
[0257] Referring to FIG. 23, when various wearable devices (2211, 2212, 2213, 2214) of FIG. 22 are linked with mobile devices (2221, 2222, 2223), the interpretation results can be displayed through the screen of the linked mobile device (2221, 2222, 2223) instead of the wearable device (2211, 2212, 2213, 2214) with a small or no screen. At this time, the linked mobile device (2221, 2222, 2223) can display the interpretation results according to the screen characteristics of each device. For example, a bar-type smart phone (2221) can display the interpretation result areas of the user and the other party separately and up and down, as in the example of 2310, and can display one result flipped over if necessary.
[0258] In the case of foldable smartphones (2222, 2223), the results of translating the user's voice can be output through the outer display, as in the examples of 2322 and 2332, and the results of translating the other party's voice can be displayed through the inner display, as in the examples of 2324 and 2334.
[0259] In this way, when a wearable device (2211, 2212, 2213, 2214) and a mobile device (2221, 2222, 2223) are linked, the wearable device (2211, 2212, 2213, 2214) receives the user's voice and, when it detects a gesture requesting translation, can display the result of translating the input user's voice in the user's interpretation result area of the mobile device (2221, 2222, 2223). At this time, the device translating the input user's voice may be a wearable device (2211, 2212, 2213, 2214) or a mobile device (2221, 2222, 2223).
[0260] In addition, when the wearable device (2211, 2212, 2213, 2214) detects a gesture requesting translation while receiving the other party's voice, the result of translating the other party's voice input before detecting the gesture requesting translation can be displayed in the area of the other party's interpretation result of the mobile device (2221, 2222, 2223). At this time, the device translating the input other party's voice may be a wearable device (2211, 2212, 2213, 2214) or a mobile device (2221, 2222, 2223).
[0261] Accordingly, even if a gesture requesting translation is input while the user or the other party is speaking, the wearable device (2211, 2212, 2213, 2214) can translate only the input voice before detecting the gesture requesting translation and display the translated result.
[0262]
[0263] FIG. 24 is a diagram illustrating an example of a gesture requesting translation in the case of a smart ring according to an embodiment of the present disclosure.
[0264] Referring to FIG. 24, the smart ring (2410) can set the gesture requesting translation as a double-tap gesture. In this case, the double-tap gesture may refer to performing the action of touching and then releasing two fingers (e.g., the thumb and index finger) twice. The smart ring (2410) can detect the corresponding gesture by detecting the number of times the two fingers touch each other using the included inertial sensor.
[0265]
[0266] According to one embodiment, a translation method using gestures may include: an operation of acquiring sensor data; an operation of determining a gesture of an electronic device and a state of the electronic device through the sensor data; an operation of determining whether a screen direction of the electronic device is directed toward a user or a counterparty through the state of the electronic device when a gesture requesting translation is detected; an operation of determining whether acquired voice data exists; an operation of translating the acquired voice data into a user's language or a counterparty's language according to the screen direction of the electronic device if the acquired voice data exists; and an operation of outputting a translated text.
[0267] In one embodiment, the gesture requesting the translation may be a gesture corresponding to voice activity detection (VAD) that identifies the end of an utterance.
[0268] According to one embodiment, if the acquired voice data exists, the operation of translating the acquired voice data into the user's language or the other party's language depending on the screen orientation of the electronic device may include: if the acquired voice data is the other party's voice and the screen orientation of the electronic device is directed toward the user, the operation of translating the acquired voice data into the user's language; and if the acquired voice data is the user's voice and the screen orientation of the electronic device is directed toward the other party, the operation of translating the acquired voice data into the other party's language.
[0269] According to one embodiment, the translation method using gestures may further include an action of acquiring the user's voice or the other party's voice according to the screen direction of the electronic device for the next translation after the action of outputting the translated text or if the acquired voice data does not exist.
[0270] According to one embodiment, the operation of acquiring the user's voice or the other party's voice according to the screen orientation of the electronic device for the next translation may include: an operation of acquiring the user's voice if the screen orientation of the electronic device is directed toward the user; and an operation of acquiring the other party's voice if the screen orientation of the electronic device is directed toward the other party.
[0271] According to one embodiment, if the screen direction of the electronic device is oriented toward the user, the operation of acquiring the user's voice may include an operation of converting the acquired user's voice into text and displaying the converted text in a direction that the user can read.
[0272] According to one embodiment, if the screen direction of the electronic device is oriented toward the other party, the operation of acquiring the other party's voice may include an operation of converting the acquired other party's voice into text and displaying the converted text in a direction that the other party can read.
[0273] According to one embodiment, the operation of outputting the translated text may include an operation of displaying the translated text corresponding to the other party's voice in a readable direction for the user when the screen direction of the electronic device is oriented toward the user; and an operation of displaying the translated text corresponding to the user's voice in a readable direction for the other party when the screen direction of the electronic device is oriented toward the other party.
[0274] According to one embodiment, the operation of outputting the translated text may include an operation of converting the translated text into an audio signal and outputting it.
[0275] According to one embodiment, the gesture requesting the translation may be at least one of a movement changing the screen orientation of the electronic device toward the user, a movement changing the screen orientation of the electronic device toward the other party, a movement bringing the electronic device closer to the user, and a movement bringing the electronic device closer to the other party.
[0276] According to one embodiment, the method includes an operation of determining whether the screen orientation of the electronic device is oriented toward the user or toward the other party, including an operation of determining whether the screen orientation of the electronic device is oriented toward the user or toward the other party as a result of determining the state of the electronic device.
[0277] According to one embodiment, a translation method using gestures may further include, prior to the translating operation, an operation of determining the language of the user and the language of the other party.
[0278] According to one embodiment, the operation of determining the user's language may include at least one of: an operation of determining a preset first language as the user's language; an operation of determining a language set in the electronic device as the user's language; an operation of determining a representative language of the region where the electronic device is located as the user's language; and an operation of executing a translation application and analyzing the initially input user voice to determine a language obtained as the user's language.
[0279] According to one embodiment, the operation of determining the language of the other party may include at least one of: an operation of determining a preset second language as the language of the other party; an operation of determining a representative language of the region where the electronic device is located as the language of the other party; and an operation of executing a translation application and analyzing the initially input voice of the other party to determine a language obtained as the language of the other party.
[0280] According to one embodiment, the operation of outputting the translated text may include: an operation of measuring a distance from the user if the translated text is the other party's voice, and an operation of measuring a distance from the other party if the translated text is the user's voice; and an operation of adjusting the size of the translated text and displaying it in consideration of the measured distance, or an operation of adjusting the volume size of a converted audio signal corresponding to the translated text and outputting it in consideration of the measured distance.
[0281] According to one embodiment, the operation of measuring the distance to the user if the translated text is the other party's voice, and measuring the distance to the other party if the translated text is the user's voice, may include measuring the distance to the user or the other party using a distance measured by a distance detection sensor, measuring the distance to the user using a Time Difference of Arrival (TDoA) method using the user's voice previously received through at least two microphones, or measuring the distance to the other party using the TDoA method using the other party's voice previously received through at least two microphones.
[0282] According to one embodiment, the electronic device may be a smart watch.
[0283] In a computer-readable recording medium, commands are stored, and when the commands are executed by one or more processors, the commands can perform the following actions: acquiring sensor data; determining a gesture of an electronic device and a state of the electronic device through the sensor data; when a gesture requesting translation is detected, determining whether a screen direction of the electronic device is toward a user or a counterparty through the state of the electronic device; determining whether acquired voice data exists; if the acquired voice data exists, translating the acquired voice data into a user's language or a counterparty's language according to the screen direction of the electronic device; and outputting a translated text.
[0284] According to one embodiment, an electronic device includes an inertial sensor for measuring inertial sensor data; one or more microphones for receiving voice; one or more processors; and a memory for storing commands, wherein the commands, when executed by the one or more processors, cause the electronic device to perform the following operations: acquiring sensor data; determining a gesture of the electronic device and a state of the electronic device based on the sensor data; when a gesture requesting translation is detected, determining whether a screen direction of the electronic device is directed toward a user or a counterparty based on the state of the electronic device; determining whether acquired voice data exists; and, if acquired voice data exists, translating the acquired voice data into a language of the user or a language of the counterparty according to a screen direction of the electronic device; and outputting a translated text.
[0285] According to one embodiment, a translation method using a gesture may include: an operation of acquiring sensor data of a wearable device from a wearable device; an operation of determining a gesture of the wearable device and a state of the wearable device through the sensor data from the wearable device; an operation of determining whether a screen direction of the wearable device is oriented toward a user or a counterparty through the state of the wearable device when a gesture requesting translation is detected from the wearable device; an operation of determining whether voice data acquired from the wearable device exists; an operation of transmitting the acquired voice data to a mobile device so as to translate the acquired voice data into a user's language or a counterparty's language according to the screen direction of the wearable device, if the acquired voice data exists; an operation of translating the acquired voice data received from the wearable device into a user's language or a counterparty's language from the mobile device; and an operation of outputting a translated text from the mobile device.
[0286]
[0287] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may store program commands, data files, data structures, etc., singly or in combination. The program commands recorded on the medium may be those specially designed and configured for the embodiment or may be those known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of the program commands include not only machine language codes generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0288] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to perform a desired operation or, independently or collectively, command the processing device. The software and / or data may be stored on any type of machine, component, physical device, virtual equipment, computer storage medium, or device, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems, and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0289] It will be appreciated that the various embodiments of the present disclosure, as described in the claims and specification, may be realized in the form of hardware, software, or a combination of hardware and software.
[0290] Such software may be stored on a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores one or more computer programs (software modules), and the one or more computer programs include computer-executable instructions that, when executed alone or collectively by one or more processors of the electronic device, cause the electronic device to perform the methods of the present disclosure.
[0291] Such software may be stored in a volatile or non-volatile storage form, for example, in the form of a storage device such as a read only memory (ROM), whether erasable or rewritable, or in the form of a memory such as a random access memory (RAM), a memory chip, device or integrated circuit, or in the form of an optically or magnetically readable medium such as a compact disk (CD), a digital versatile disc (DVD), a magnetic disk or magnetic tape or the like. It will be appreciated that such storage devices and storage media are various embodiments of non-transitory machine-readable storage devices suitable for storing a computer program or computer programs comprising instructions that, when executed, implement various embodiments of the present disclosure. Accordingly, various embodiments provide a program comprising code for implementing an apparatus or method according to any claim of the present disclosure and a non-transitory machine-readable storage medium storing such a program.
[0292] While the present disclosure has been described and illustrated with reference to various embodiments thereof, it will be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope defined by the appended claims and their equivalents.
Claims
1. One or more non-transitory computer-readable storage media containing computer-executable instructions that, when individually or collectively executed by one or more processors of an electronic device, cause the electronic device to perform the following operations: The above action is, The act of acquiring sensor data; An operation of determining a gesture of the electronic device and a state of the electronic device through the sensor data; When a gesture requesting translation is detected, an action of checking whether the screen direction of the electronic device is directed toward the user or toward the other party through the status of the electronic device; An action to check whether acquired voice data exists; If the acquired voice data exists, an operation of translating the acquired voice data into the user's language or the other party's language according to the screen orientation of the electronic device; and Action to output translated text One or more non-transitory computer-readable storage media containing.
2. In paragraph 1, The gesture requesting the above translation is, A gesture corresponding to voice activity detection (VAD) that confirms the end of speech. One or more non-transitory computer-readable storage media.
3. In paragraph 1, If the above-mentioned acquired voice data exists, the operation of translating the above-mentioned acquired voice data into the user's language or the other party's language according to the screen orientation of the electronic device is as follows: If the acquired voice data is the voice of the other party and the screen direction of the electronic device is directed toward the user, an operation of translating the acquired voice data into the user's language; and If the acquired voice data is the user's voice and the screen direction of the electronic device is directed toward the other party, an operation of translating the acquired voice data into the other party's language. One or more non-transitory computer-readable storage media containing:
4. In paragraph 1, After the action of outputting the translated text or if the acquired voice data does not exist, an action of acquiring the user's voice or the other party's voice according to the screen direction of the electronic device for the next translation. One or more non-transitory computer-readable storage media further comprising:
5. In paragraph 1, The operation of acquiring the user's voice or the other party's voice according to the screen direction of the electronic device for the above-mentioned next translation is as follows: If the screen direction of the electronic device is directed toward the user, an operation of acquiring the user's voice; and If the screen direction of the electronic device is directed toward the other party, an operation of acquiring the other party's voice One or more non-transitory computer-readable storage media containing:
6. In paragraph 1, If the screen direction of the electronic device is directed toward the user, the operation of acquiring the user's voice is: An action of converting the acquired user's voice into text and displaying the converted text in a direction that the user can read. One or more non-transitory computer-readable storage media containing:
7. In paragraph 1, If the screen direction of the electronic device is directed toward the other party, the operation of acquiring the other party's voice is: An action of converting the acquired voice of the other party into text and displaying the converted text in a direction that the other party can read. One or more non-transitory computer-readable storage media containing:
8. In paragraph 1, The action of outputting the translated text above is: If the screen direction of the electronic device is directed toward the user, an action of displaying translated text corresponding to the other party's voice in a direction that the user can read; and If the screen direction of the electronic device is directed toward the other party, an action of displaying translated text corresponding to the user's voice in a direction that the other party can read. One or more non-transitory computer-readable storage media containing:
9. In paragraph 1, The action of outputting the translated text above is: An action to convert the translated text above into an audio signal and output it. One or more non-transitory computer-readable storage media containing:
10. In paragraph 1, The gesture requesting the above translation is, A movement in which the screen orientation of the electronic device changes to face the user; A movement in which the screen orientation of the electronic device changes to face the other party; a movement that brings said electronic device closer to said user, or A movement of bringing said electronic device closer to said opponent; One or more non-transitory computer-readable storage media, at least one of which is:
11. In paragraph 1, The action of checking whether the screen direction of the above electronic device is directed toward the user or toward the other party is as follows: An operation of determining whether the screen of the electronic device is tilted toward the user or toward the other party as a result of judging the state of the electronic device. One or more non-transitory computer-readable storage media containing:
12. In paragraph 1, Before the above translation action, Action to determine the language of the user and the language of the other party Including more, The action of determining the language of the above user is: An action to determine the preset first language as the language of the user; An action to determine the language set in the electronic device as the language of the user; An action to determine the representative language of the region where the electronic device is located as the language of the user; or An action in which a translation application is executed, the language obtained by analyzing the initially input user voice is determined as the user's language. Contains at least one of the following: The action of determining the language of the other party is, An action of determining a preset second language as the language of the other party; An action to determine the representative language of the region where the electronic device is located as the language of the other party; or An action in which a translation application is executed, the language obtained by analyzing the first inputted voice of the other party is determined as the language of the other party. One or more non-transitory computer-readable storage media containing at least one of:
13. In paragraph 1, The action of outputting the translated text above is: An operation of measuring the distance from the user if the translated text is the other party's voice, and measuring the distance from the other party if the translated text is the user's voice; and An operation of adjusting the size of the translated text and displaying it in consideration of the measured distance, or adjusting the volume size of the converted audio signal corresponding to the translated text and outputting it in consideration of the measured distance. One or more non-transitory computer-readable storage media containing:
14. In paragraph 13, The operation of measuring the distance from the user if the translated text is the other party's voice, and measuring the distance from the other party if the translated text is the user's voice, Measure the distance from the user or the other party using the distance measured through the distance detection sensor, or Measure the distance to the user using the Time Difference of Arrival (TDoA) method using the user's voice previously received through at least two microphones, or Measuring the distance to the counterparty using the TDoA method using the counterparty's voice previously received through at least two microphones One or more non-transitory computer-readable storage media.
15. In the translation method using gestures, The act of acquiring sensor data; An action of judging a gesture of an electronic device and a state of the electronic device through the above sensor data; When a gesture requesting translation is detected, an action of checking whether the screen direction of the electronic device is directed toward the user or toward the other party through the status of the electronic device; An action to check whether acquired voice data exists; If the acquired voice data exists, an operation of translating the acquired voice data into the user's language or the other party's language according to the screen orientation of the electronic device; and Action to output translated text How to include.
Citation Information
Patent Citations
A method and apparatus for translating foreign language using user's gesture, and mobile device using the same
KR1020110039660A
Device and method for voice translation
KR1020170112713A
Vehicle seat fixing device
KR102443869B1
Wearable device and translation system
US20160267075A1
Wearable translation device
US20160283469A1