Electronic device for providing sign language service using graphic object, operating method thereof, and storage medium

The electronic device synchronizes avatar motions with speaking speed and uses distinct graphic objects for each speaker to address the challenge of providing high-quality sign language services, ensuring clear and synchronized sign language interpretation.

WO2025216622A1PCT designated stage Publication Date: 2025-10-16SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/099764
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-11
Filing Date
2025-03-13
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing electronic devices struggle to provide high-quality sign language services by synchronizing avatar motions with the speaking speed of multiple speakers, leading to disconnection and difficulty in identifying the current speaker.

Method used

An electronic device identifies individual speakers through voice recognition, adjusts motion playback times for sign language sentences to match speaking durations, and uses distinct graphic objects for each speaker to ensure temporal synchronization and clear identification.

Benefits of technology

The solution enables hearing-impaired individuals to easily recognize spoken content by providing synchronized and differentiated sign language services, enhancing the clarity and quality of the sign language interpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025099764_16102025_PF_FP_ABST
    Figure KR2025099764_16102025_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment, electronic devices (101, 201) comprise one or more processors (120, 320) and memories (130, 330) storing instructions. The instructions, when executed by the one or more processors, may be configured to cause the electronic devices to: identify a plurality of speakers on the basis of speech data included in content; identify a sign language sentence corresponding to an utterance related to a first speaker among the plurality of speakers; identify a duration time of the utterance and a motion time for the sign language sentence; adjust motion playback times for each of words included in the sign language sentence such that the motion time for the sign language sentence corresponds to the duration time of the utterance; and express the words included in the sign language sentence by using a graphic object related to the first speaker according to the adjusted motion playback times.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device for providing sign language service using graphic objects, its operation method and storage medium

[0001] One embodiment disclosed in this document relates to an electronic device for providing a sign language service using a graphic object, a method of operating the same, and a storage medium.

[0002] The variety of services and additional features offered through electronic devices such as smartphones is steadily increasing. To enhance the utility of these devices and satisfy the diverse needs of users, telecommunications service providers and electronic device manufacturers are competitively developing electronic devices that offer diverse features and differentiate themselves from competitors. Consequently, the various functions offered through electronic devices are also becoming increasingly sophisticated. Recently, various types of intelligence services for electronic devices have been offered, and among these intelligence services, avatars can be used to provide diverse services to users.

[0003] Electronic devices can utilize avatars to provide sign language services, chat services, game application services, or virtual education services. For example, if sign language is provided through a speaker's avatar, and the avatar is synchronized temporally to the speaker's speaking speed, hearing-impaired people can easily and clearly perceive the content of the speech, enabling them to receive high-quality sign language services.

[0004] According to one embodiment, an electronic device includes at least one processor and a memory storing instructions, wherein the instructions, when executed by the at least one processor, are configured to cause the electronic device to identify a plurality of speakers based on voice data included in content.

[0005] According to one embodiment, the instructions may be configured to cause the electronic device to identify a sign language sentence corresponding to an utterance associated with a first speaker among the plurality of speakers.

[0006] According to one embodiment, the instructions may be configured to cause the electronic device to identify the duration of the utterance and the motion time for the sign language sentence.

[0007] According to one embodiment, the instructions may be configured to cause the electronic device to adjust motion playback times for each word included in the sign language sentence such that the motion time for the sign language sentence corresponds to the duration of the utterance.

[0008] According to one embodiment, the instructions may be configured to cause the electronic device to represent words included in the sign language sentence using graphic objects associated with the first speaker according to preset motion playback times.

[0009] According to one embodiment, a method for providing a sign language service using a graphic object in an electronic device may include an operation of identifying a plurality of speakers based on voice data included in content.

[0010] In one embodiment, the method may include an operation of identifying a sign language sentence corresponding to an utterance associated with a first speaker among the plurality of speakers.

[0011] In one embodiment, the method may include an operation of identifying a duration of the utterance and a motion time for the sign language sentence.

[0012] According to one embodiment, the method may include adjusting motion playback times for each word included in the sign language sentence so that the motion time for the sign language sentence corresponds to the duration of the utterance.

[0013] According to one embodiment, the method may include an operation of expressing words included in the sign language sentence using a graphic object associated with the first speaker according to the adjusted motion playback times.

[0014] According to one embodiment, a non-transitory storage medium storing instructions, wherein the instructions, when executed by at least one processor of an electronic device, are configured to cause the electronic device to perform at least one operation, wherein the at least one operation may include an operation of identifying a plurality of speakers based on voice data included in content.

[0015] In one embodiment, the at least one action may include an action of identifying a sign language sentence corresponding to an utterance associated with a first speaker among the plurality of speakers.

[0016] In one embodiment, the at least one action may include an action of identifying a duration of the utterance and a motion time for the sign language sentence.

[0017] According to one embodiment, the at least one action may include adjusting motion playback times for each of the words included in the sign language sentence so that the motion time for the sign language sentence corresponds to the duration of the utterance.

[0018] According to one embodiment, the at least one action may include an action of expressing words included in the sign language sentence using a graphic object associated with the first speaker according to the adjusted motion playback times.

[0019] A computer-readable non-transitory recording medium according to one embodiment of the present disclosure may store at least one command and / or instructions that, when executed, cause an electronic device to perform the method or operation of the electronic device described above.

[0020] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.

[0021] FIG. 2 is a diagram illustrating a method for separating a voice signal within content according to one embodiment.

[0022] Figure 3 is an internal block diagram of an electronic device for providing sign language service according to one embodiment.

[0023] FIG. 4a is a diagram illustrating a method for recognizing a speaker's voice according to one embodiment.

[0024] FIG. 4b is a diagram illustrating a method for identifying a speaker when a speaker list does not exist according to one embodiment.

[0025] FIG. 4c is a diagram illustrating a method for identifying a speaker when a speaker list exists according to one embodiment but does not match previously registered speaker data.

[0026] FIG. 4d is a diagram illustrating a method for identifying a speaker when a speaker list exists according to one embodiment but matches previously registered speaker data.

[0027] Figure 5 is a flowchart illustrating the operation of an electronic device for providing sign language service according to one embodiment.

[0028] FIG. 6 is a diagram for explaining a sign language conversion operation according to one embodiment.

[0029] FIG. 7 is a diagram illustrating a method for calculating sign language playback time according to one embodiment.

[0030] FIG. 8 is a diagram for explaining a method for calculating sign language reproduction time when the number of words in a spoken sentence and the number of words in a sign language sentence are the same according to one embodiment.

[0031] Figure 9 is a diagram showing an adjusted sign language playback time according to one embodiment.

[0032] FIG. 10 is a diagram for explaining a method for calculating sign language reproduction time when the number of words in a spoken sentence and the number of words in a sign language sentence are different according to one embodiment.

[0033] Figure 11 is an example screen diagram providing a sign language service using a graphic object for each speaker according to one embodiment.

[0034] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0035] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to one embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0036] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0037] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0038] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0039] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0040] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0041] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0042] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0043] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0044] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0045] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0046] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0047] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0048] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0049] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0050] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0051] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0052] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0053] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0054] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0055] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0056] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0057] In the detailed description below, reference numerals in the drawings may be used interchangeably or omitted for components that can be easily understood through the preceding embodiments, and their detailed descriptions may also be omitted. The electronic device (101) according to one embodiment disclosed in this document may be implemented by selectively combining components of different embodiments, and components of one embodiment may be replaced by components of another embodiment. For example, it should be noted that the present invention is not limited to specific drawings or embodiments.

[0058] Before describing the present invention, a method for providing sign language services using avatars will be described. Avatars may be graphic objects, graphic icons, or animated characters representing users of electronic devices.

[0059] Sign language services using avatars are services that interpret (or convert, or translate) a speaker's speech into sign language through an avatar representing the speaker. For example, a sign language service might use an avatar's hand motions to express spoken content while outputting (or displaying) content.

[0060] In the case of sign language, the sentence structure is different from that of general Korean, so a process of translating spoken sentences into sign language is necessary, and a transformer model is used to translate spoken sentences into sign language.

[0061] Sign language is a spatial language, utilizing not only the hands but also the upper body, such as the face. Therefore, even words representing the same vocabulary in spoken and signed language can have different utterance lengths. For example, when providing sign language through a speaker's avatar, the motions used to express the sign language may cause the motion playback time to differ from the utterance length. Therefore, it is necessary to temporally synchronize the avatar to the speaker's speaking rate to express the sign language. Furthermore, when multiple speakers speak simultaneously, their words can be mixed and translated into sign language, requiring sign language services to be differentiated for each speaker.

[0062] In one embodiment, an electronic device, an operating method thereof, and a storage medium for providing a temporally synchronized sign language service using a graphic object corresponding to a speaker in response to the speech speed of each speaker may be provided.

[0063] To this end, in one embodiment, an electronic device (101) can identify speakers based on their voices and map avatars corresponding to the identified speakers. The electronic device (101) can provide sign language services by temporally synchronizing the corresponding avatars according to the speech speed of each speaker. By doing so, a hearing-impaired person can easily and clearly recognize the content of the speech, and by providing sign language services by matching different avatars to each speaker, the current speaker can be easily recognized, thereby providing high-quality sign language services.

[0064] Before explaining the present invention, a method for separating a voice signal within content used in the present invention will be described with reference to FIG. 2. FIG. 2 is a drawing for explaining a method for separating a voice signal within content according to one embodiment.

[0065] The types of content provided by electronic devices (101) are becoming more diverse. For example, the content may be video content filmed by a user, or video content such as dramas or movies downloaded or streamed from at least one server (e.g., content provider). In the following description, content includes visual images and auditory sounds, and may be content including images and voices containing people, for example. In addition, video content such as dramas or movies may include subtitles (or subtitle data). Here, the image may refer to a moving video, and the voice may be described as audio or sound.

[0066] Referring to FIG. 2, when playing content on an electronic device (101), voice signals (205) containing the voices of multiple speakers included in the content may be output through a speaker. The voice signals (205) may be separated into individual signals based on sound separation technology (210). The electronic device (101) may identify a speaker by extracting voice features for each of the separated voice signals and match the identified speaker (220).

[0067] Specifically, the electronic device (101) can identify (or separate, classify) each of a plurality of voice signals within the content and match the identified voice signals to each speaker. In one embodiment, the electronic device (101) can separate voice signals within the content and, for example, identify speakers by referencing pre-stored speaker data using the separated voice signals through learning such as deep learning.

[0068] According to one embodiment, the electronic device (101) can identify speakers by separating voice signals for each speaker in the content, and can match an avatar corresponding to the identified speaker.

[0069] Meanwhile, after identifying a speaker and matching an avatar, in order to express sign language corresponding to the utterance through the avatar, it is necessary to express sign language through the avatar in a temporally synchronized manner corresponding to the speaker's utterance duration (or utterance length, utterance speed).

[0070] In one embodiment, the motion playback time for each word included in the sign language sentence corresponding to the utterance for each speaker can be adjusted (or synchronized) (230) to control the voice utterance speed and the playback speed of the sign language motion using the avatar, thereby providing a clear and accurate sign language service.

[0071] Figure 3 is an internal block diagram of an electronic device for providing sign language service according to one embodiment.

[0072] Referring to FIG. 3, an electronic device (201) according to an embodiment (e.g., the electronic device (101) of FIG. 1) may include at least one processor (320) (e.g., the processor (120) of FIG. 1) and a memory (530) (e.g., the memory (130) of FIG. 1). According to an embodiment, the electronic device (201) may include a display (360) (e.g., the display module (160) of FIG. 1) and / or a communication circuit (390) (e.g., the communication module (190) of FIG. 1). Here, not all components illustrated in FIG. 3 are essential components of the electronic device (201), and the electronic device (201) may be implemented by more or fewer components than the components illustrated in FIG. 3.

[0073] According to one embodiment, the memory (330) may store a control program for controlling the electronic device (201), a UI related to an application provided by the manufacturer or downloaded from an external source, images for providing the UI, user information, documents, databases, or related data.

[0074] According to one embodiment, the memory (330) may store a voice recognition module (or voice recognition engine). The voice recognition module may include an AI-based speaker recognition model. Examples of speaker recognition models include a chirp model or a deep speaker model, but the types thereof are not limited thereto.

[0075] Meanwhile, voice recognition can be performed by the electronic device (201) itself, but this is only an example, and the electronic device (201) can also perform voice recognition through a server (208) connected to the electronic device (201). For example, the processor (320) can transmit voice according to speech to the server (208) through a communication circuit (390), and receive the voice recognition result from the server (208) through the communication circuit (390).

[0076] According to one embodiment, the processor (320) may control the overall operation of an electronic device (201) having a function for providing a sign language service using a graphic object. For example, the processor (320) may be implemented as an application processor (AP).

[0077] According to one embodiment, the processor (320) can translate a speaker's speech into sign language. For example, when playing content, the processor (320) can convert voice data included in the content into text, which is character data, through speech recognition (STT) technology. The processor (320) can translate the converted text into sign language through an artificial intelligence-based sign language translation model. The processor (320) can play a motion (or action) corresponding to the translated sign language through a graphic object of the corresponding speaker.

[0078] In one embodiment, in order to express sign language by temporally synchronizing motions through graphic objects in response to the speaking speed of the speaker, the processor (320) may adjust the playback time of motions corresponding to sign language in response to the speaking speed. The processor (320) may adjust the playback time (or playback length, playback speed) of motions for each word included in the sign language sentence according to the time (or speech length, speech speed) of speaking each word included in the speech so that the speech duration (or total speech time) and the motion time of the translated sign language (or sign language sentence) correspond (or match). Accordingly, the processor (320) may provide a sign language service that suits the speaking speed of the speaker by expressing each word included in the sign language through the motion of the graphic object in accordance with the adjusted motion playback times. By controlling the motion playback time of the graphic object for sign language by considering the utterance time of each word included in the utterance, the utterance speed and the graphic object during content playback can be temporally synchronized to express sign language, thereby preventing the sense of disconnection due to the time difference.

[0079] According to one embodiment, the operation for adjusting the playback time of the motion of a graphic object for sign language by considering the utterance time of each word included in the utterance may be performed by the processor (320), but may also be performed by components such as those illustrated in FIG. 4A. FIG. 4A is a diagram illustrating a method for recognizing a speaker's voice according to one embodiment.

[0080] Referring to FIG. 4A, the electronic device (201) may include a speaker voice recognition unit (410), a voice length analysis unit (420), a sign language conversion unit (430), and / or a sign language image output unit (440). In addition, the speaker voice recognition unit (410) may include a voice analysis unit (412), a speaker classification result unit (414), and / or a speaker data unit (416).

[0081] Referring to FIG. 4A, the speaker voice recognition unit (410) (or processor (320)) can extract voice features based on voice data included in the content (405) and identify at least one speaker based on the extracted voice features. For example, the content (405) can be divided into video data and voice data, and among them, the voice data can be input to the speaker voice recognition unit (410). In addition, video content such as dramas or movies can include subtitles (or subtitle data), and the subtitle data can be input to the voice length analysis unit (420).

[0082] The speaker voice recognition unit (410) (e.g., voice analysis unit (412)) can analyze the voice data using the voice recognition module stored in the memory (330). For example, the voice analysis unit (412) can extract voice features of the voice data and transmit the extracted voice features to the speaker classification result unit (414) so ​​that the speaker can be identified. Here, the extracted voice features can include unique voice features in addition to male voice features, female voice features, and child voice features.

[0083] The speaker classification result section (414) may refer to the speaker data section (416) in the memory (330) for speaker identification based on the extracted voice features. Here, the speaker data section (416) serves to store and manage a speaker list, and the speaker list may include voice features for multiple speakers and / or identification information (e.g., voice ID) of speakers corresponding to the voice features.

[0084] According to one embodiment, the speaker voice recognition unit (410) (or processor (320)) may store and register voice features of the plurality of speakers and / or speaker identification information corresponding to the voice features in the memory (330). The speaker classification result unit (414) may distinguish whether the extracted voice feature is a speaker registered in the speaker list or a new speaker.

[0085] For example, the speaker classification result unit (414) (or processor (320)) can obtain speaker identification information based on a result that matches a certain percentage or more by comparing extracted voice features with pre-registered voice features in the speaker list with reference to the speaker list.

[0086] Referring to FIG. 4b, which is a drawing for explaining a method for identifying a speaker when a speaker list does not exist according to an embodiment, when there is no speaker list stored in the speaker data unit (416), the speaker classification result unit (414) can generate identification information (e.g., voice ID) corresponding to the voice features extracted for registration as a new speaker. Here, the speaker classification result unit (414) can generate identification information (e.g., avatar information) of a graphic object to be mapped to the new speaker. Accordingly, the speaker data unit (416) can designate and manage a graphic object that replaces the new speaker.

[0087] Meanwhile, referring to FIG. 4c, which is a drawing for explaining a method for identifying a speaker when a speaker list exists according to an embodiment but does not match the previously registered speaker data, the speaker classification result unit (414) may compare the extracted voice features with the previously registered voice features (e.g., speaker data 1, speaker data 2, …) in the speaker list, and if none of the previously registered voice features matches the extracted voice features, for example, if the comparison results do not match by a certain percentage or more, it may determine that there is no registered speaker. Accordingly, the speaker classification result unit (414) may generate identification information (e.g., voice ID) corresponding to the extracted voice features in order to register a speaker corresponding to the extracted voice features as a new speaker, and may also generate identification information of a graphic object for the new speaker. Accordingly, the speaker data unit (416) may designate and manage a graphic object for the newly registered speaker.

[0088] Referring to FIG. 4d, which is a drawing for explaining a method of identifying a speaker when a speaker list exists according to an embodiment and matches pre-registered speaker data, the speaker classification result unit (414) compares the extracted voice features with pre-registered voice features in the speaker list, and if there is a pre-registered voice feature that matches the extracted voice feature, for example, if the comparison result matches by a certain percentage or more, the speaker corresponding to the extracted voice feature can be identified. Accordingly, the speaker classification result unit (414) can confirm a designated graphic object for the registered speaker by using the speaker's identification information (e.g., speaker ID).

[0089] As described above, the processor (320) can classify multiple speakers recognized from voice data into registered speakers and unregistered speakers by referring to the speaker list of the speaker data unit (416). If the speaker's identification information is not obtained, the processor (320) can classify the speaker as an unregistered speaker. On the other hand, if the speaker's identification information is obtained, the processor (320) can classify the speaker as a registered speaker. The processor (320) can map (or designate) a designated (or set) graphic object to a registered speaker, and map a new graphic object to an unregistered speaker.

[0090] The processor (320) may, by referring to the speaker list, utilize a graphic object set for a registered speaker if one exists. Furthermore, if a speaker is registered in the speaker list but does not have a set graphic object, the processor (320) may designate a new graphic object for the registered speaker, and may also designate another new graphic object for a speaker not registered in the speaker list.

[0091] In this way, the processor (320) can set different graphic objects for each speaker, whether the speaker is registered or not. The processor (320) can output (or play, display) a video and a sign language video according to content playback on the display (360), and can configure the sign language video as a part of the content playback screen. For example, the processor (320) can provide sign language interpretation through graphic objects that are set differently for each speaker on the content playback screen. The processor (320) can output (or display) the graphic objects on the display (360) so that it is easy to distinguish who the speaker of the voice currently being interpreted in sign language is.

[0092] The processor (320) can reproduce a motion (or action) corresponding to the translated sign language through a graphic object of the speaker. Here, the graphic object may include an avatar representing a character representing the speaker.

[0093] By expressing sign language through graphic objects set differently for each speaker in this way, users can easily recognize the speaker currently speaking.

[0094] The processor (320) can express the motion of a graphic object representing the speaker as translated sign language movements. For example, the processor (320) can translate predefined hand gestures and other body movements for sign language sentences into actions of an avatar.

[0095] Meanwhile, in the above, an example was described in which a speaker is identified through voice recognition of voice data included in the content when the content is played, but a voice spoken by a user is received through a microphone (not shown) of an electronic device (201), and voice recognition is performed on the user's voice to obtain text for the voice and translate it into sign language, and the method for obtaining a voice used for converting it into sign language is not limited thereto.

[0096] As described above, the processor (320) may perform voice recognition on voice output during content playback or on voice received via a microphone. According to one embodiment, the voice recognition processing for speech may partially include automatic speech recognition (ASR) processing. For example, the speaker's voice may be converted into text using an ASR module (or ASR engine), and the voice according to the user's speech may be converted into text based on the voice recognition results.

[0097] According to one embodiment, the voice recognition processing may be processed in a voice recognition module (or voice recognition engine) stored in the memory (330) of the electronic device (201), or in a server (208) (e.g., an intelligent server and / or a service server). Accordingly, the processor (320) may transmit the voice spoken by the user to the server (208) or transmit the voice recognition result to the server (208).

[0098] Meanwhile, in order to adjust the playback time of the motion of the graphic object for sign language by considering the utterance time of each word included in the utterance, the voice length analysis unit (420) can obtain input text through voice recognition for the utterance, and can analyze (or calculate) the total time (or utterance duration) of the utterance sentence and the length (or time) of each word included in the utterance sentence from the obtained input text. In addition, video content such as dramas or movies can include subtitles (or subtitle data), and the voice length analysis unit (420) can analyze not only the total time of the utterance sentence but also the length of each word in the utterance sentence based on the subtitle data.

[0099] For example, if the first speaker utters 'I'll go in first today' for 01:13 to 01:17 seconds, and the second speaker utters 'Okay, see you tomorrow' for 01:19 to 01:22 seconds, the voice length analysis unit (420) can check the duration of the first speaker's utterance (e.g., 00:04 seconds) and the length (or time, speed) of each of the four words 'today', 'I', 'first', and 'I'll go in' within the utterance sentence for the first speaker. In addition, the voice length analysis unit (420) can check the duration of the second speaker's utterance (e.g., 00:03 seconds) and the length of each of the three words 'Okay', 'tomorrow', and 'see you' within the utterance sentence for the second speaker.

[0100] According to one embodiment, the voice length analysis unit (420) can provide information about the time of the spoken sentence and the length of each word in the spoken sentence for each speaker to the sign language conversion unit (430).

[0101] According to one embodiment, the sign language conversion unit (430) can translate (or convert) a spoken sentence into sign language based on the speaker's identification information transmitted from the speaker voice recognition unit (410) and the information on the speaking time and the length of each word in the spoken sentence transmitted from the voice length analysis unit (420). The sign language conversion unit (430) can adjust the playback time (or playback speed, playback length) of the motion of the graphic object differently for each word when playing back a motion corresponding to the translated sign language through the speaker's graphic object.

[0102] According to one embodiment, the sign language conversion unit (430) may convert an input text corresponding to an utterance of a first speaker into a sign language sentence based on an artificial intelligence model. For example, a transformer model may be used to translate the utterance sentence into sign language. The converted sign language sentence is composed of a plurality of words (or a word list), and the sign language conversion unit (430) may refer to a sign language motion-word dictionary unit (e.g., the sign language motion-word dictionary unit (434) of FIG. 6) that stores the relationship between sign language motions and words, and may match (or search, identify) a motion corresponding to each word in the sign language sentence. On the other hand, for words that do not exist in the sign language motion-word dictionary unit, the sign language conversion unit (430) may match them to motions that express them literally.

[0103] The sign language conversion unit (430) can check the utterance duration and the motion time for the sign language sentence, and adjust the motion time for the sign language sentence or the motion playback times of each word included in the sign language sentence so that the utterance duration and the motion time for the sign language sentence that are required to make the sign language into a motion correspond (or become the same).

[0104] According to one embodiment, the sign language conversion unit (430) can check the number of words included in the spoken sentence and the number of words in the sign language sentence, and adjust the motion playback times of each word included in the sign language sentence based on the check result.

[0105] If the number of words included in the spoken sentence is the same as the number of words in the signed sentence, the sign language conversion unit (430) may calculate a ratio between the duration of the spoken sentence and the motion time for the signed sentence and adjust the motion time for the signed sentence to increase or decrease by the calculated ratio. According to one embodiment, when adjusting the motion time for the signed sentence to increase or decrease, the sign language conversion unit (430) may adjust the motion playback times for each word included in the signed sentence by the calculated ratio.

[0106] On the other hand, if the number of words included in the spoken sentence is different from the number of words in the sign language sentence, the sign language conversion unit (430) may adjust the first motion playback time for words included in the sign language sentence that match the words included in the spoken sentence, and may adjust the second motion playback time for words included in the sign language sentence that do not match the words included in the spoken sentence. For example, the motion playback time may be adjusted at a certain ratio for words included in the sign language sentence that match the words included in the spoken sentence, and may be adjusted at a second motion playback time that is different from the certain ratio, that is, different from the first motion playback time, for words included in the sign language sentence that do not match the words included in the spoken sentence. For example, if the number of words included in the spoken sentence is greater than the number of words in the sign language sentence, the sign language conversion unit (430) may adjust the second motion playback time so that words included in the unmatched sign language sentence are operated as sign language motions within the remaining time excluding the first motion playback time from the total time of the spoken sentence.

[0107] Meanwhile, the sign language conversion unit (430) can identify a sign language sentence corresponding to an utterance related to a second speaker, and identify a duration of the utterance related to the second speaker and a motion time for the sign language sentence corresponding to the utterance related to the second speaker. The sign language conversion unit (430) can adjust motion playback times for each word included in the sign language sentence corresponding to the utterance related to the second speaker so that the motion time for the sign language sentence corresponding to the utterance related to the second speaker corresponds to the duration of the utterance related to the second speaker, and can express the words included in the sign language sentence corresponding to the utterance related to the second speaker using a graphic object related to the second speaker according to the adjusted motion playback times.

[0108] The sign language video output unit (440) can output (or display, play) a sign language video using a graphic object corresponding to each speaker according to the motion playback speed adjusted in the sign language conversion unit (430).

[0109] Meanwhile, at least some of the above-described operations of the electronic device (201) may be performed by a server (208) (e.g., an intelligent server and / or a service server).

[0110] According to one embodiment, an electronic device (101, 201) includes at least one processor (120, 32) and a memory (130, 330) storing instructions, which, when executed by the at least one processor, may be configured to cause the electronic device to identify a plurality of speakers based on voice data included in content.

[0111] According to one embodiment, the instructions may be configured to cause the electronic device to identify a sign language sentence corresponding to an utterance associated with a first speaker among the plurality of speakers.

[0112] According to one embodiment, the instructions may be configured to cause the electronic device to identify the duration of the utterance and the motion time for the sign language sentence.

[0113] According to one embodiment, the instructions may be configured to cause the electronic device to adjust motion playback times for each word included in the sign language sentence such that the motion time for the sign language sentence corresponds to the duration of the utterance.

[0114] According to one embodiment, the instructions may be configured to cause the electronic device to represent words included in the sign language sentence using graphic objects associated with the first speaker according to preset motion playback times.

[0115] According to one embodiment, the instructions may be configured to cause the electronic device to obtain an input text for the utterance, identify the number of words included in the input text and the number of words in the sign language sentence, and adjust motion playback times for each of the words included in the sign language sentence based on the results of identifying the number of words included in the input text and the number of words in the sign language sentence.

[0116] According to one embodiment, the instructions may be configured to cause the electronic device to calculate a ratio between the duration of the utterance and the motion time for the sign language sentence when the number of words included in the input text is the same as the number of words in the sign language sentence, and to adjust motion playback times for each of the words included in the sign language sentence using the calculated ratio.

[0117] According to one embodiment, the instructions may be configured to cause the electronic device to, when the number of words included in the input text is different from the number of words in the sign language sentence, identify whether words included in the input text match words included in the sign language sentence, adjust a motion playback time for words in the sign language sentence that match words included in the input text to a first motion playback time, and adjust a motion playback time for words in the sign language sentence that do not match words included in the input text to a second motion playback time.

[0118] According to one embodiment, the instructions may be configured to cause the electronic device to adjust the second motion playback time according to the remaining motion playback time excluding the first motion playback time from the duration of the utterance when the number of words in the sign language sentence is greater than the number of words included in the input text.

[0119] According to one embodiment, the instructions may be configured to cause the electronic device to identify the time of each word included in the input text and the duration of the utterance based on at least one of subtitle data included in the content and input text for the utterance.

[0120] According to one embodiment, the instructions may be configured to cause the electronic device to extract voice features from voice data included in the content and identify the plurality of speakers based on the extracted voice features.

[0121] According to one embodiment, the instructions may be configured to cause the electronic device to classify the plurality of speakers into registered speakers and unregistered speakers using a pre-stored plurality of speaker lists based on the extracted voice features, map a designated graphic object to the registered speakers, and map a new graphic object to the unregistered speakers.

[0122] According to one embodiment, the electronic device further comprises a display, and the instructions are configured to cause the electronic device to map a graphical object associated with each of the identified plurality of speakers and display each of the mapped graphical objects on the display, wherein the graphical object may include an avatar.

[0123] According to one embodiment, the instructions may be configured to cause the electronic device to identify a sign language sentence corresponding to an utterance related to a second speaker among the plurality of speakers, identify a duration of the utterance related to the second speaker and a motion time for the sign language sentence corresponding to the utterance related to the second speaker, adjust motion playback times for each of words included in the sign language sentence corresponding to the utterance related to the second speaker so that the motion time for the sign language sentence corresponding to the utterance related to the second speaker corresponds to the duration of the utterance related to the second speaker, and express words included in the sign language sentence corresponding to the utterance related to the second speaker using a graphic object related to the second speaker according to the adjusted motion playback times.

[0124] FIG. 5 is a flowchart illustrating an operation of an electronic device for providing a sign language service according to an embodiment. Referring to FIG. 5, the operation method may include operations 505 to 525. Each operation of the operation method of FIG. 5 may be performed by at least one of an electronic device (e.g., the electronic device (101) of FIG. 1, the electronic device (201) of FIG. 3, or at least one processor of the electronic device (e.g., the processor (120) of FIG. 1, the processor (320) of FIG. 3)). In an embodiment, at least one of operations 505 to 525 may be omitted, the order of some operations may be changed, or another operation may be added.

[0125] In operation 505, the electronic device (201) can identify multiple speakers based on voice data included in the content.

[0126] According to one embodiment, the electronic device (201) can extract voice features from voice data included in the content and identify the plurality of speakers based on the extracted voice features.

[0127] According to one embodiment, the electronic device (201) can classify the plurality of speakers into registered speakers and unregistered speakers using a list of multiple speakers stored in advance based on the extracted voice features.

[0128] According to one embodiment, the electronic device (201) can map a designated graphic object for the registered speaker and map a new graphic object for the unregistered speaker.

[0129] In operation 510, the electronic device (201) can identify a sign language sentence corresponding to an utterance related to a first speaker among the plurality of speakers.

[0130] In operation 515, the electronic device (201) can identify the duration of the utterance and the motion time for the sign language sentence.

[0131] According to one embodiment, the electronic device (201) can identify the time of each word included in the input text and the duration of the utterance based on at least one of the subtitle data included in the content and the input text for the utterance.

[0132] In operation 520, the electronic device (201) can adjust the motion playback times for each word included in the sign language sentence so that the motion time for the sign language sentence corresponds to the duration of the utterance.

[0133] According to one embodiment, the operation of adjusting motion playback times for each word included in the sign language sentence may include the operation of obtaining an input text for the utterance, the operation of identifying the number of words included in the input text and the number of words in the sign language sentence, and the operation of adjusting motion playback times for each word included in the sign language sentence based on the result of identifying the number of words included in the input text and the number of words in the sign language sentence.

[0134] According to one embodiment, the operation of adjusting motion playback times for each word included in the sign language sentence may include, when the number of words included in the input text and the number of words in the sign language sentence are the same, the operation of calculating a ratio between the duration of the utterance and the motion time for the sign language sentence, and the operation of adjusting motion playback times for each word included in the sign language sentence using the calculated ratio.

[0135] According to one embodiment, the operation of adjusting motion playback times for each word included in the sign language sentence may include, when the number of words included in the input text and the number of words in the sign language sentence are the same, the operation of calculating a ratio between the duration of the utterance and the motion time for the sign language sentence, and the operation of adjusting motion playback times for each word included in the sign language sentence using the calculated ratio.

[0136] According to one embodiment, the operation of adjusting motion playback times for each word included in the sign language sentence may include, when the number of words included in the input text and the number of words in the sign language sentence are different, an operation of identifying whether words included in the input text match words included in the sign language sentence, an operation of adjusting motion playback times for words in the sign language sentence that match words included in the input text to a first motion playback time, and an operation of adjusting motion playback times for words in the sign language sentence that do not match words included in the input text to a second motion playback time.

[0137] According to one embodiment, the operation of adjusting the motion playback time for a word of the sign language sentence to a second motion playback time may include an operation of adjusting the second motion playback time according to the remaining motion playback time excluding the first motion playback time from the duration of the utterance when the number of words of the sign language sentence is greater than the number of words included in the input text.

[0138] In operation 525, the electronic device (201) can express words included in the sign language sentence using graphic objects related to the first speaker according to the adjusted motion reproduction times.

[0139] According to one embodiment, the electronic device (201) can map a graphic object associated with each of the identified plurality of speakers. The electronic device (201) can display each of the mapped graphic objects, and the graphic objects can include avatars.

[0140] FIG. 6 is a diagram for explaining a sign language conversion operation according to an embodiment. To help understand the description of FIG. 6, reference will be made to FIGS. 7 to 10. FIG. 7 is a diagram for explaining a method for calculating a sign language reproduction time according to an embodiment, FIG. 8 is a diagram for explaining a method for calculating a sign language reproduction time when the number of words in an utterance sentence and the number of words in a sign language sentence are the same according to an embodiment, FIG. 9 is a diagram for illustrating an adjusted sign language reproduction time according to an embodiment, and FIG. 10 is a diagram for explaining a method for calculating a sign language reproduction time when the number of words in an utterance sentence and the number of words in a sign language sentence are different according to an embodiment.

[0141] Referring to FIG. 6, according to one embodiment, the sign language conversion unit (430) may include a sign language sentence conversion unit (432), a sign language motion-word dictionary unit (434), a motion generation unit (436), and / or an avatar sign language generation unit (438).

[0142] Referring to FIG. 6, the sign language sentence conversion unit (432) can convert an input text corresponding to a first speaker's utterance into a sign language sentence. FIG. 6 exemplifies a case where the sign language sentence conversion unit (432) converts text-type subtitle data into a sign language sentence, but it is also possible to convert voice data for the first speaker's utterance into text and then convert the converted text into a sign language sentence. In addition, the sign language sentence conversion unit (432) can also use both the converted text and the text-type subtitle data to convert into sign language sentences.

[0143] The converted sign language sentence is composed of multiple words (or word lists), and the sign language motion-word dictionary unit (434) can match (or search, identify) a motion corresponding to each word in the sign language sentence by referring to data that stores the relationship between sign language motions and words, and can provide the matching result to the motion generation unit (436). On the other hand, for words that do not exist in the sign language motion-word dictionary unit, the result of matching a motion that directly expresses the word can be provided to the motion generation unit (436).

[0144] The motion generation unit (436) may receive a word list, i.e., each word included in a sign language sentence, from the sign language sentence conversion unit (432), and may receive a motion matching result, i.e., a motion list, from the sign language motion-word dictionary unit (434). In addition, the motion generation unit (436) may receive information on the duration of speech and the time of each word included in the speech sentence from the voice length analysis unit (420). Here, the motion generation unit (436) may map the motion and playback speed for each word in the form of a table as illustrated in FIG. 6 for each speaker at each time of speech.

[0145] For example, referring to Figure 7, when the original sentence (700) uttered by the first speaker is 'Hello? It seems like the weather is really clear today', the total time of the utterance sentence (or utterance duration) (720) may be the sum of the times for each word (710) in the utterance sentence. Here, the total time of the utterance sentence is t sentence It can be said that the motion generation unit (436) can compare the number of words included in the original sentence (700) with the number of words included in the sign language sentence (or sign language sentence) (730).

[0146] According to one embodiment, the number of words included in a spoken sentence and the number of words included in a sign language sentence may be compared, and a method of adjusting motion playback times for each word included in a sign language sentence may vary depending on whether the number of each word is the same.

[0147] First, when the number of words included in the spoken sentence and the number of words included in the signed sentence are the same, the motion generation unit (436) calculates a ratio between the duration of the spoken sentence (or the total time of the spoken sentence) and the motion time for the signed sentence (or the motion time of the entire signed sentence), and can adjust the motion playback times for each word included in the signed sentence at a certain ratio using the calculated ratio.

[0148] Specifically, the motion generation unit (436) generates the time of the matched word (e.g., word utterance time) (t word )(740) and the motion playback time (t), which is the time required to move the corresponding word as a motion of sign language. motion )(750) using the ratio (ω) of the converted playback time (or speed, length) (t play ) can be obtained. According to one embodiment, the motion generation unit (436) converts the play time (or speed, length) (t) after conversion so that the sign language motion is played for the total sentence time (720), that is, so that the sign language play time (770) becomes the same as the total sentence time (720). play ) can be adjusted.

[0149] For example, referring to Fig. 8(a), the sign language conversion unit (430) (or motion generation unit (436)) generates the entire original sentence utterance time (t sentence) (805) The entire motion time of the sign language sentence (t) motion )(810) is larger, the entire utterance time of the original sentence (t sentence) The entire motion time (t') of the sign language sentence to correspond to (or be identical to) (805) motion )(815) can be adjusted to decrease.

[0150] In contrast, referring to Fig. 8(b), the sign language conversion unit (430) (or motion generation unit (436)) outputs the entire original sentence utterance time (t sentence) (820) The entire motion time of the sign language sentence (t) motion )(825) is smaller, the entire utterance time of the original sentence (t sentence) The entire motion time (t') of the sign language sentence to correspond to (or be identical to) (820) motion )(830) can be adjusted to increase.

[0151] Accordingly, the motion generation unit (436) generates the playback times (or speed, length) (t) after conversion, as shown in FIG. 9. play1 + t play2 + t play3 + t play4 ) By adjusting the motion playback time of each word to correspond to the total sentence time (720), the sign language playback time (770) can be increased or decreased by a certain amount of time (900). To this end, the motion generation unit (436) can adjust the motion playback time based on the following mathematical expression 1.

[0152]

[0153] Based on the above mathematical expression 1, the motion generation unit (436) calculates the ratio (ω) between the duration of the utterance and the motion time for the sign language sentence. total ) and the calculated ratio (ω total ) can be used to adjust the motion playback times for each word included in a sign language sentence. For example, the calculated ratio (ω total )(or coefficient) can be multiplied by each motion playback speed to calculate the motion playback time. If the entire original sentence utterance time (t sentence) = 10s, and the entire motion time of the sign language sentence (t motion ) = 5s, if each of the 5 motion images is 1s, the playback speed coefficient (ω) total) = 2. Therefore, for each image, the corresponding playback speed coefficient (ω total = 2) is applied, the playback length of each motion can be 1s x 2 = 2s. At this time, the motion generation unit (436) can add movement time for continuous movement between each motion. Here, the added time can be calculated using the average speed of the two motions and the distance between the start and end points of the motions.

[0154] According to one embodiment, by matching the speaking time of each word spoken by the speaker for each sign language motion, expression according to the speaking time of the speaker is possible, thereby providing a high-quality sign language service.

[0155] Meanwhile, when the number of words included in the spoken sentence and the number of words included in the sign language sentence are different, the motion generation unit (436) can calculate a ratio between the duration of the utterance (or the total time of the spoken sentence) and the motion time for the sign language sentence (or the motion time of the entire sign language sentence), and then adjust the motion playback times for each word at different ratios according to the matching result between the words included in the spoken sentence and the words included in the sign language sentence. For example, the motion generation unit (436) can adjust the motion playback time for the words of the sign language sentence that match the words included in the spoken sentence to the first motion playback time, and adjust the motion playback time for the words of the sign language sentence that do not match the words included in the spoken sentence to the second motion playback time.

[0156] Referring to Figure 10, the total sentence time (1020) (T sentence ) = 5s, and the time per word in the utterance sentence (1040) (e.g., word utterance time) (t word1, t word2, t word3, t word4 ) each t word1 = 1.5s, t word2 = 1s, t word3 = 1s, tword4 = 1.5s, and the motion time of each word in the sign language sentence (t motion1, t motion2, t motion3, t motion4 )(1st motion playback time)(1050) each t motion1 = 3s, t motion2 = 3s, t motion3 = 3s, t motion4-1 = 2s, t motion4-2 = When 4s is said, t is the word in the sign language sentence that matches the word in the spoken sentence. word1, t word2, t word3 The playback speed coefficient and playback time can be calculated based on the following mathematical formula 2.

[0157]

[0158] Referring to the above mathematical expression 2, t, a word in a sign language sentence word1, t word2, t word3 The playback speed coefficients for are ω1 = 0.5, ω2 = 0.666, ω3 = 0.666, and the playback time (second motion playback time) (1055) is t' motion1 = 1.5s, t' motion2 = 1s, t' motion3 = can be obtained with 1s.

[0159] Accordingly, the motion generation unit (436) can calculate the final playback speed coefficient and playback time of the words in the sign language sentence that match the words in the spoken sentence based on the following mathematical expression 3.

[0160]

[0161] Referring to the above mathematical expression 3, the final playback speed coefficient ω' of the word in the sign language sentence matching the word in the spoken sentence is 5 / ((1.5 + 1 + 1) + (5 - (1.5 + 1 + 1)) / 2) = 1.1764, and the adjusted motion playback time (1060) is t play1 = 1.275s, t play2= 0.85s, t play3 = can be obtained as 0.85s.

[0162] On the other hand, the total time length of words in a sign language sentence that do not match words in an utterance sentence can be calculated based on the following mathematical formula 4.

[0163]

[0164] Referring to the above mathematical expression 4, the total time length (T) of the words in the unmatched sign language sentence unmacthed ) is the total sentence time (1020)(T sentence )(e.g. 5s) Time lengths of matched words (e.g. t play1, t play2, t play3 ) can be obtained based on the remaining time excluding (or subtracting) the time. For example, the total time length of the words in the unmatched sign language sentence (T unmacthed ) is the total sentence time (1020)(T sentence )(eg 5s) - t play1 (e.g. 1.275s) - t play2 (e.g. 0.85s) - t play3 (e.g. 0.85s) = 2.025.

[0165] The motion generation unit (436) calculates the total time length (T) of words in the unmatched sign language sentences. unmacthed ), the playback speed coefficient (ω4) of the unmatched word can be obtained as 2.025 / (2 + 4) = 0.3375. Accordingly, the playback time (or playback length, speed) of the unmatched word is t play4-1 = 2s x 0.3375 = 0.675s, t play4-2 = 4s x 0.3375 = 1.35s. Accordingly, the motion generation unit (436) adjusts the motion playback time of each word included in the sign language sentence, so the total motion playback time (1070) of the sign language sentence is, for example, t play1 (e.g. 1.275s) + t play2 (e.g. 0.85s) +tplay3 (e.g. 0.85s) + t play4-1 (e.g. 0.675s) + t play4-2 (e.g. 1.35s) = 5s can be adjusted to be equal to the total time of the utterance sentence.

[0166] As described above, the motion generation unit (436) can check the motion time for the sign language sentence and adjust the motion time for the sign language sentence or the motion playback times of each word included in the sign language sentence so that the duration of the utterance and the motion time for the sign language sentence required to operate the sign language as a motion correspond (or become the same).

[0167] According to one embodiment, the motion generation unit (436) can check the number of words included in the spoken sentence and the number of words in the sign language sentence, and adjust the motion playback times of each word included in the sign language sentence based on the check result.

[0168] If the number of words included in the spoken sentence is the same as the number of words in the signed sentence, the motion generation unit (436) may calculate a ratio between the duration of the spoken sentence and the motion time for the signed sentence and adjust the motion time for the signed sentence to increase or decrease by the calculated ratio. According to one embodiment, the sign language conversion unit (430) may adjust the motion playback times for each word included in the signed sentence by the calculated ratio when adjusting the motion time for the signed sentence to increase or decrease.

[0169] On the other hand, if the number of words included in the spoken sentence is different from the number of words in the sign language sentence, the motion generation unit (436) may adjust the first motion playback time for words included in the sign language sentence that match the words included in the spoken sentence, and may adjust the second motion playback time for words included in the sign language sentence that do not match the words included in the spoken sentence. For example, the motion playback time may be adjusted at a certain ratio for words included in the sign language sentence that match the words included in the spoken sentence, and may be adjusted at a second motion playback time that is different from the certain ratio, that is, different from the first motion playback time, for words included in the sign language sentence that do not match the words included in the spoken sentence. For example, if the number of words included in the spoken sentence is greater than the number of words in the sign language sentence, the motion generation unit (436) may adjust the second motion playback time so that words included in the unmatched sign language sentence are operated as sign language motions within the remaining time excluding the first motion playback time from the entire time of the spoken sentence.

[0170] Accordingly, the avatar sign language generation unit (438) can express words included in a sign language sentence using graphic objects related to the speaker according to motion playback times adjusted for each word.

[0171] Figure 11 is an example screen diagram providing a sign language service using a graphic object for each speaker according to one embodiment.

[0172] Referring to FIG. 11, the electronic device (201) can display graphic objects (e.g., avatars) mapped to each speaker on the screen while playing content (1100). As illustrated in FIG. 11, by using different graphic objects (1105, 1110, 1115, 1120, 1125) for each speaker to display (or play) sign language, a hearing-impaired person can clearly recognize what kind of conversation each speaker is having.

[0173] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0174] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0175] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0176] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0177] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0178] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0179] According to one embodiment, in a non-transitory storage medium storing instructions, the instructions, when executed by at least one processor (120, 320) of an electronic device (101, 201), are set to cause the electronic device to perform at least one operation, wherein the at least one operation may include: an operation of identifying a plurality of speakers based on voice data included in content; an operation of identifying a sign language sentence corresponding to an utterance related to a first speaker among the plurality of speakers; an operation of identifying a duration of the utterance and a motion time for the sign language sentence; an operation of adjusting motion playback times for each of words included in the sign language sentence so that the motion time for the sign language sentence corresponds to the duration of the utterance; and an operation of expressing words included in the sign language sentence using a graphic object related to the first speaker according to the adjusted motion playback times.

Claims

1. In the electronic device (101, 201), At least one processor (120, 320); and Contains a memory (130, 330) for storing instructions, The above instructions, when executed by the at least one processor, cause the electronic device to: Identify multiple speakers based on voice data contained in the content, Identifying a sign language sentence corresponding to an utterance related to a first speaker among the above plurality of speakers, Identify the duration of the above utterance and the motion time for the above sign language sentence, Adjust the motion playback times for each word included in the sign language sentence so that the motion time for the sign language sentence corresponds to the duration of the utterance, An electronic device configured to express words included in the sign language sentence using graphic objects related to the first speaker according to the adjusted motion playback times.

2. In the first paragraph, the instructions cause the electronic device to: Obtain the input text for the above utterance, Identify the number of words contained in the input text and the number of words in the sign language sentence, An electronic device configured to adjust motion playback times for each word included in the sign language sentence based on the result of identifying the number of words included in the input text and the number of words in the sign language sentence.

3. In the first or second paragraph, the instructions cause the electronic device to: If the number of words included in the input text is the same as the number of words in the sign language sentence, the ratio between the duration of the utterance and the motion time for the sign language sentence is calculated, An electronic device configured to adjust motion playback times for each word included in the sign language sentence using the above-described ratio.

4. In any one of paragraphs 1 to 3, the instructions cause the electronic device to: If the number of words included in the input text is different from the number of words in the sign language sentence, identify whether the words included in the input text and the words included in the sign language sentence match, Adjust the motion playback time for the words of the sign language sentence matching the words included in the input text to the first motion playback time, An electronic device set to adjust the motion playback time for words in the sign language sentence that do not match words included in the input text to a second motion playback time.

5. In any one of paragraphs 1 to 4, the instructions cause the electronic device to: If the number of words in the sign language sentence is greater than the number of words contained in the input text above, An electronic device set to adjust the second motion playback time according to the remaining motion playback time excluding the first motion playback time from the duration of the utterance.

6. In any one of paragraphs 1 to 5, the instructions cause the electronic device to: An electronic device configured to identify the time of each word included in the input text and the duration of the utterance based on at least one of subtitle data included in the content and input text for the utterance.

7. In any one of paragraphs 1 to 6, the instructions cause the electronic device to: Extract voice features from the voice data included in the above content, An electronic device configured to identify the plurality of speakers based on the extracted voice features.

8. In any one of paragraphs 1 to 7, the instructions cause the electronic device to: Based on the extracted voice features, the plurality of speakers are classified into registered speakers and unregistered speakers using a list of multiple speakers stored in advance. For the above registered speaker, map the specified graphic object, An electronic device configured to map new graphic objects to the above unregistered speakers.

9. In any one of paragraphs 1 to 8, further comprising a display, The above instructions cause the electronic device to: Mapping a graphic object associated with each of the plurality of speakers identified above, An electronic device configured to display each of the above-mentioned mapped graphic objects on the display, wherein the graphic objects include avatars.

10. In any one of paragraphs 1 to 9, the instructions cause the electronic device to: Identifying a sign language sentence corresponding to an utterance related to a second speaker among the above multiple speakers, Identify the duration of the utterance associated with the second speaker and the motion time for the sign language sentence corresponding to the utterance associated with the second speaker, Adjust the motion playback times for each word included in the sign language sentence corresponding to the utterance related to the second speaker so that the motion time for the sign language sentence corresponding to the utterance related to the second speaker corresponds to the duration of the utterance related to the second speaker, An electronic device configured to express words included in sign language sentences corresponding to utterances related to the second speaker using graphic objects related to the second speaker according to the above-described controlled motion playback times.

11. A method for providing sign language service using graphic objects in an electronic device, An action to identify multiple speakers based on voice data contained in the content; An action of identifying a sign language sentence corresponding to an utterance related to a first speaker among the plurality of speakers; An action of identifying the duration of the above utterance and the motion time for the above sign language sentence; An operation of adjusting the motion playback times for each word included in the sign language sentence so that the motion time for the sign language sentence corresponds to the duration of the utterance; and A method for providing a sign language service using a graphic object, including an action of expressing words included in the sign language sentence using a graphic object related to the first speaker according to the adjusted motion playback times.

12. In paragraph 11, the operation of adjusting the motion playback times for each word included in the sign language sentence is as follows: An action of obtaining input text for the above utterance; An operation of identifying the number of words contained in the input text and the number of words in the sign language sentence; and A method for providing a sign language service using a graphic object, comprising an operation of adjusting motion playback times for each word included in the sign language sentence based on the result of identifying the number of words included in the input text and the number of words in the sign language sentence.

13. In the 11th or 12th paragraph, the operation of adjusting the motion playback times for each word included in the sign language sentence is as follows: An operation of calculating a ratio between the duration of the utterance and the motion time for the sign language sentence when the number of words included in the input text is the same as the number of words in the sign language sentence; and A method for providing a sign language service using a graphic object, including an operation of adjusting motion playback times for each word included in the sign language sentence using the above-described calculated ratio.

14. In any one of paragraphs 11 to 13, the operation of adjusting motion playback times for each word included in the sign language sentence is as follows: An operation of identifying whether the words included in the input text and the words included in the sign language sentence match when the number of words included in the input text and the number of words in the sign language sentence are different; An operation of adjusting the motion playback time for a word of the sign language sentence matching a word included in the input text to a first motion playback time; and A method for providing a sign language service using a graphic object, comprising an action of adjusting the motion playback time for a word of the sign language sentence that does not match a word included in the input text to a second motion playback time.

15. In a non-transitory storage medium storing instructions, the instructions are configured to cause the electronic device (101, 201) to perform at least one operation when executed by at least one processor (120, 320), wherein the at least one operation is: An action to identify multiple speakers based on voice data contained in the content; An action of identifying a sign language sentence corresponding to an utterance related to a first speaker among the plurality of speakers; An action of identifying the duration of the above utterance and the motion time for the above sign language sentence; An operation of adjusting the motion playback times for each word included in the sign language sentence so that the motion time for the sign language sentence corresponds to the duration of the utterance; and A storage medium including an action of expressing words included in the sign language sentence using a graphic object related to the first speaker according to the adjusted motion playback times.

Citation Information

Patent Citations

  • Sign language translation device and program

    JP2019124901A

  • Sign language word time length calculation device and program thereof, and sign language cg video generation device and program thereof

    JP2023173001A

  • The sign language providing system using auto-transformed voice recognition data

    KR1020130065064A

  • Knee protector

    KR102431035B1

  • Trampoline fitness equipment

    KR102550737B1