Method for providing sign language video and electronic device performing same

The electronic device addresses the challenge of expressing difficult words in sign language by generating a composite video with a sign language video and sub-image, effectively conveying meaning and reducing delays.

WO2025249853A1PCT designated stage Publication Date: 2025-12-04SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/007089
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-20
Filing Date
2025-05-26
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Converting speech or text into sign language is challenging for words like proper nouns, neologisms, and technical terms, as they are difficult to express and delay communication due to separate expression of consonants and vowels, making it hard to convey clear meaning and context.

Method used

An electronic device determines a first set of words for generating a sign language video and a second set of words for generating a sub-image, using analysis information to create a composite video that includes both sign language video and sub-image to convey meaning effectively.

Benefits of technology

The solution enhances communication by clearly conveying meaning through sign language videos with sub-images for difficult words, reducing animation delays and resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025007089_04122025_PF_FP_ABST
    Figure KR2025007089_04122025_PF_FP_ABST
Patent Text Reader

Abstract

A method for providing a sign language video, performed by an electronic device, may comprise the operations of: determining analysis information for an original text; determining, on the basis of the analysis information, a first word set for generating a sign language video and a second word set for generating sub-images of the sign language video, from among words included in the original text; generating a sign language video on the basis of one or more words included in the first word set; generating sub-images of the sign language video on the basis of one or more words included in the second word set; generating a composite video including the sign language video and the sub-images; and displaying the composite video via a display of the electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

Method for providing sign language video and electronic device for performing the same

[0001] Below, a technology for providing sign language video based on text is disclosed.

[0002] Sign language is a standalone language, like Korean or English. When converting speech or text into sign language, words and expressions difficult to express in sign language, such as proper nouns, neologisms, and technical terms, can be represented through finger language. When finger language is used to represent words, the consonants and vowels are expressed separately, which delays communication and makes it difficult to clearly convey meaning and context.

[0003] Technologies are being developed to provide video content, including sign language animation, to assist deaf people in conveying information and communicating. Sign language video content can enhance the effectiveness of its message through auxiliary means such as virtual avatars, text, or images.

[0004] According to one embodiment, a method performed by an electronic device may include an operation of determining analysis information for an original text. The method may include an operation of determining, based on the analysis information, a first set of words for generating a sign language video and a second set of words for generating a sub-image of the sign language video from among words included in the original text. The method may include an operation of generating a sign language video based on one or more words included in the first set of words. The method may include an operation of generating a sub-image of the sign language video based on one or more words included in the second set of words. The method may include an operation of generating a composite video including the sign language video and the sub-image. The method may include an operation of displaying the composite video through a display of the electronic device.

[0005] According to one embodiment, a non-transitory computer-readable recording medium may store one or more programs including instructions. When the instructions are individually or collectively executed by at least one processor of an electronic device, the instructions may cause the electronic device to: determine analysis information for an original text. When the instructions are individually or collectively executed by the at least one processor, the electronic device may: determine a first word set for generating a sign language video and a second word set for generating a sub-image of the sign language video from among words included in the original text based on the analysis information. When the instructions are individually or collectively executed by the at least one processor, the electronic device may: generate a sign language video based on one or more words included in the first word set. When the instructions are individually or collectively executed by the at least one processor, the electronic device may: generate a sub-image of the sign language video based on one or more words included in the second word set. When the above commands are individually or collectively executed by the at least one processor, the electronic device may be caused to: generate a composite video including the sign language video and the sub-image. When the above commands are individually or collectively executed by the at least one processor, the electronic device may be caused to: display the composite video through a display of the electronic device.

[0006] According to one embodiment, an electronic device includes at least one processor including processing circuitry, and a memory including one or more storage media storing instructions, wherein the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to: determine analysis information for an original text. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to: determine a first set of words for generating a sign language video from among words included in the original text and a second set of words for generating a sub-image of the sign language video based on the analysis information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to: generate a sign language video based on one or more words included in the first set of words. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to: generate a sub-image of the sign language video based on one or more words included in the second set of words. When the above commands are individually or collectively executed by the at least one processor, the electronic device may be caused to: generate a composite video including the sign language video and the sub-image. When the above commands are individually or collectively executed by the at least one processor, the electronic device may be caused to: display the composite video through a display of the electronic device.

[0007] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.

[0008] FIG. 2 is a diagram illustrating an artificial intelligence system according to one embodiment.

[0009] Figure 3 is a schematic diagram of an electronic device according to one embodiment.

[0010] Figure 4 is a flowchart of a method for providing sign language video according to one embodiment.

[0011] FIG. 5 is a flowchart of a method for determining a first word set and a second word set according to one embodiment.

[0012] FIG. 6 is a diagram illustrating a method for determining a first word set and a second word set according to an example.

[0013] Figure 7 is a flowchart of a method for generating sign language video according to one embodiment.

[0014] FIG. 8 is a diagram illustrating a method for generating a sign language video according to one embodiment.

[0015] FIG. 9 is a flowchart of a method for generating a sub-image according to one embodiment.

[0016] FIG. 10 is a flowchart of a method for generating a sub-image according to one embodiment.

[0017] FIG. 11 and FIG. 12 are drawings each illustrating a method for generating a sub-image according to an example.

[0018] FIG. 13 is a flowchart of a method for generating a synthetic video according to one embodiment.

[0019] FIG. 14 and FIG. 15 are drawings each illustrating a method for generating a synthetic video according to an example.

[0020] Hereinafter, an embodiment of the present document may be described with reference to the attached drawings.

[0021] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.

[0022] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0023] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0024] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0025] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0026] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0027] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0028] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0029] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0030] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0031] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0032] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0033] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0034] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0035] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0036] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0037] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0038] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0039] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0040] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0041] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0042] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0043] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service by itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0044] FIG. 2 is a diagram illustrating an artificial intelligence system according to one embodiment.

[0045] In an artificial intelligence system (hereinafter, “system”) (200), a User Query / Response Interface (User Query / Response Interface) (210) can receive user input. The user input can be any type of input, such as natural language, image, audio, and / or video. Additionally, context information can be transmitted together when the user input is transmitted. The context information can include various side information related to the time when the user input is input into the system (200). For example, the context information can include information about the application currently being used by the user or information about the user’s location. Additionally, the user input can be a mixed type of input of the above-described natural language, image, audio, video, and / or context information. Additionally, the user input can include non-natural language input, such as selecting a menu.

[0046] A user query / response interface (210) can provide output from a generative artificial intelligence system to a user. The output may include a natural language-based response and / or specific content. The output may also include an action requested by the user.

[0047] The AI ​​framework (220) can receive user input. Based on the user input (e.g., the user's query), the AI ​​framework (220) can coordinate and control one or more components necessary to perform an action corresponding to the user's intent.

[0048] User input received from the user query / response interface (210) can be transmitted to a prompt design component (221). The prompt design component (221) can be used to generate a prompt suitable as input to a generative model (e.g., a large language model (LLM) and / or a large multimodal model (LMM)) based on the user input.

[0049] The prompt design component (221) may be an AI component that uses a machine learning algorithm or a neural network. The prompt design component (221) may generate improved prompts through learning over time. The prompt design component (221) may access a knowledge repository (230) to generate prompts based on user input. The knowledge repository (230) may include user preference data, a prompt library, and / or prompt examples. The prompt design component (223) may provide the generated prompts to a generative model (e.g., an LLM and / or an LMM).

[0050] The APIs / Plugins management component (223) can communicate with external information sources based on requests for additional information when user input is transmitted to the generative model.

[0051] The APIs / Plugins management component (223) can establish a communication channel for communication with the outside of the system (200) via the API. The APIs / Plugins management component (223) can enable access to various data sources via the communication channel. The acquired information can be used to generate prompts by the prompt design component (221) along with user input, or can be used as input for the generative model (250).

[0052] The APIs / Plugins management component (223) can request a final action via an API when the final action in response to user input, rather than an intermediate action, must be performed by an application or service.

[0053] The refiner component (225) can fine-tune the output of the generative model (250). For example, the refiner component (225) can determine the relevance (e.g., score) between the output (e.g., content) of the generative model and the user input. For example, the refiner component (225) can determine whether the output contains biased information (e.g., selective information). For example, the refiner component (225) can determine whether the output contains harmful information (e.g., violent content or profanity).

[0054] The refinement component (225) can determine the degree of matching (e.g., score) between the output of the generative model (250) and the user input (e.g., the intent of the user input). If the refinement component (225) determines that the output of the generative model (250) does not correspond to the user input, the refinement component (225) can modify the output so that it corresponds to the user input.

[0055] The refinement component (225) can provide hints (e.g., hints for prompt generation) to the user so that the user can obtain information that matches the user's intention from the generative model (250).

[0056] A generative model (250) may refer to an artificial intelligence neural network that generates new data (e.g., text, images, audio, or video) based on user input (e.g., user utterances). The generative model (250) may include an image generation model and / or a language generation model.

[0057] Image generation models may include generative adversarial networks (GANs) and / or variational autoencoders (VAEs). An example of an image generation model is a diffusion-based generative model with a VAE and transformer architecture.

[0058] A language generation model (e.g., ChatGPT) can be a model trained to generate statistically most appropriate output based on input. A language generation model can include an LLM. An LLM can identify various types of input, such as text, images, audio (e.g., speech), and / or video, and generate new data corresponding to the input.

[0059] Figure 3 is a schematic diagram of an electronic device according to one embodiment.

[0060] According to one embodiment, the electronic device (300) (e.g., the electronic device (101) of FIG. 1) may be a device such as a mobile terminal (e.g., a smartphone, a tablet, a laptop) or a stationary terminal (e.g., a personal computer (PC)). According to one embodiment, the electronic device (300) may be implemented in the form of a wearable electronic device (e.g., smart glasses, a head mounted display).

[0061] According to one embodiment, the electronic device (300) may include at least a part of the configuration of the electronic device (101) of FIG. 1. The electronic device (300) may include a processor (310) including processing circuitry (e.g., the processor (120) of FIG. 1). The processor (310) may include at least one processor. The electronic device (300) may include a memory (320) including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1). The memory (320) may include a text acquisition module (330), a text analysis module (340), a word set determination module (350), a prompt generation module (360), a sign language video generation module (370), a sub-image generation module (380), and a synthesis module (390). Each of the text acquisition module (330), the text analysis module (340), the word set determination module (350), the prompt generation module (360), the sign language video generation module (370), the sub-image generation module (380), and the synthesis module (390) may be an operation, a function, a process, or a software based on at least a part of the instructions stored in the memory (320). When the instructions stored in the memory (320) are individually or collectively executed by at least one processor of the processor (310), the electronic device (300) may be caused to perform at least a part of the sign language video providing method of the present disclosure (e.g., the operation, function, process, or software of the text acquisition module (330), the text analysis module (340), the word set determination module (350), the prompt generation module (360), the sign language video generation module (370), the sub-image generation module (380), or the synthesis module (390)).

[0062] The electronic device (300) can obtain original data for generating a sign language video. According to one embodiment, the electronic device (300) can obtain original text as original data through a text acquisition module (330). According to one embodiment, the electronic device (300) can obtain original video data or original voice data as original data. The text acquisition module (330) can include a speech recognition module (or, a speech-to-text (SST) module). The electronic device (300) can obtain original text based on the original video data or original voice data through the voice recognition module.

[0063] The electronic device (300) can analyze and preprocess the original text through the text analysis module (340). The electronic device (300) can determine analysis information about the original text through the text analysis module (340). Based on the analysis information, the electronic device (300) can convert the original text into text suitable for the word order and grammar of sign language through the text analysis module (340).

[0064] The electronic device (300) can determine a first word set for generating a sign language video and a second word set for generating a sub-image of the sign language video from among words included in the original text based on analysis information through the word set determination module (350). The electronic device (300) can determine a first word set and a second word set from among words included in a text converted to suit the word order and grammar of sign language based on analysis information through the word set determination module (350). A method for determining the first word set and the second word set is described in detail with reference to FIG. 5.

[0065] The electronic device (300) can generate a sign language video based on one or more words included in a first word set for generating a sign language video through a sign language video generation module (370). A method for generating a sign language video is described in detail with reference to FIG. 7.

[0066] The electronic device (300) can generate a sub-image based on one or more words included in a second word set for generating a sub-image of a sign language video through a sub-image generation module (380).

[0067] The electronic device (300) can determine a prompt using one or more words included in the second word set via the prompt generation module (360). Based on the prompt, the electronic device (300) can generate a sub-image using a generative artificial intelligence (AI) model (e.g., the generative model (250) of FIG. 2).

[0068] According to one embodiment, the electronic device (300) may include a generative AI model, which is an on-device model capable of generating sub-images without communication with an external electronic device (e.g., a server). For example, the sub-image generation module (380) may include the generative AI model. The electronic device (300) may input a prompt to the generative AI model and obtain a sub-image output by the generative AI model.

[0069] According to one embodiment, the electronic device (300) may transmit a prompt to an external server through communication with the external server including a generative AI model. The external server may generate a sub-image corresponding to the prompt using the generative model. The electronic device (300) may receive the sub-image from the external server. For example, the electronic device (300) may transmit a prompt to the external server including the generative AI model through the sub-image generation module (380) and receive a sub-image corresponding to the prompt from the external server.

[0070] The method of generating a sub-image is described in detail with reference to FIGS. 9 and 10.

[0071] The electronic device (300) can generate a synthetic video including a sign language video and a sub-image through a synthesis module (390). A method for generating a synthetic video is described in detail with reference to FIG. 13.

[0072] Figure 4 is a flowchart of a method for providing sign language video according to one embodiment.

[0073] According to one embodiment, the operations 410 to 450 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (300) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1 or the processor (310) of FIG. 3) and a memory (e.g., the memory (130) of FIG. 1 or the memory (320) of FIG. 3).

[0074] According to one embodiment, an electronic device can obtain source data for generating a sign language video. For example, the electronic device can obtain original text as the source data. For example, the electronic device can obtain original video data or original voice data as the source data. The electronic device can obtain the original text based on the original video data or the original voice data. The electronic device can obtain the original text by converting voice data or original voice data included in the original video data into text.

[0075] At operation 410, the electronic device can determine analysis information about the original text.

[0076] According to one embodiment, the analysis information for the original text may include at least one of the sentence types of the original text, sentence components of words included in the original text, or parts of speech of words included in the original text.

[0077] The sentence type of the original text can be one of the following: single sentence, compound sentence, or complex sentence.

[0078] A simple sentence is a sentence that contains only one independent clause.

[0079] A compound sentence is a sentence in which two or more main clauses are connected by a coordinating conjunction (e.g., and, but, or, so). Each main clause in a compound sentence can exist as an independent sentence, that is, as a simple sentence.

[0080] A complex sentence is a sentence that consists of a main clause and one or more dependent clauses connected by subordinating conjunctions (e.g., that, while, when, before, until, because, if). The main clause of a complex sentence can exist as an independent sentence (simple sentence), but the dependent clauses cannot exist as independent sentences.

[0081] Analysis information about the original text may include information about the main clause and / or subordinate clause of the original text.

[0082] The sentence components of the words included in the original text can be one of the following: subject, predicate, object, complement, modifier, adverbial, or independent element.

[0083] The parts of speech of words in the original text can be noun, pronoun, numeral, verb, adjective, adverb, determiner, interjection, or particle (or postposition).

[0084] Depending on the language of the source text, there may be differences in determining analytical information. For example, in agglutinative languages ​​like Korean, affixes can be added to roots or stems to change the function or meaning of words, and particles are attached to nouns and pronouns. In contrast, languages ​​like English and Chinese, which have little or no word morphological inflection, primarily convey grammatical meaning through grammatical words independent of word order, and prepositions are used instead of particles to indicate grammatical relationships between nouns. In Korean, adjectives are functionally similar to verbs and can be used alone or as predicates. Adjectives and determiners are used to modify nouns. Determiners modify nouns, pronouns, and substantives, which include numerals. English does not have a clear sentence element or part-of-speech category corresponding to the Korean "determinant" or "determiner," but it does have modifiers (e.g., adjectives, adverbs, prepositional phrases, participle phrases), determiners (e.g., articles, possessive adjectives, demonstrative adjectives, quantifiers), and adjectives that modify nouns, similar to subjects, objects, and complements. Furthermore, English sentences mainly follow the subject-verb-object structure, while Korean sentences mainly follow the subject-object-verb structure, and the subject can be omitted. In Korean, predicates are generally distinguished from objects and complements, whereas in English, sentences are divided into subjects and predicates, and the predicate can include not only the verb but also the objects, complements, or adverbs that the verb takes. In this specification, unless otherwise stated, the explanation is based on the case where the original text is in Korean. If the original text is in a language other than Korean, such as English, the ‘predicate’ can be understood as a ‘verb’ and the ‘adjective’ can be understood as a ‘modifier’ or ‘qualifier.’

[0085] In operation 420, the electronic device can determine a first set of words for generating a sign language video and a second set of words for generating a sub-image of the sign language video from among words included in the original text based on the analysis information.

[0086] According to one embodiment, the electronic device can pre-store sign language dictionary information and sign language animation information for generating sign language videos.

[0087] Sign language dictionary information may include words that can be expressed in sign language. The words included in the sign language dictionary information can be categorized into representative entries mapped to sign language animations, entries that can be expressed with the same sign language animation as the representative entry (homonymous words), or compound words formed by combining representative entries.

[0088] Sign language dictionary information may include associations between representative headwords and homonyms, as well as associations between representative headwords and compound words. For example, among the words "추정" (chujeong), "약열" (yakpyol), "강탈" (gangtal), "빼다" (bakda), and "거로채다" (geochaeda), the sign language dictionary information may include associations between the representative headword ("거로채다") and homonyms ("추정" (chujeong), "약열" (yakpyol), "강탈" (gangtal), and "빼다") mapped to sign language animation.

[0089] Sign language animation information may include sign language animations mapped to (or corresponding to) multiple representative headings among the words included in the sign language dictionary information. The sign language animation may represent a video or compressed data of a video in which a virtual avatar expresses the sign language movements of the representative headings.

[0090] According to one embodiment, the electronic device can generate a sub-image visualizing words included in the original text that are not included in the sign language dictionary information or that correspond to a predetermined type. Words not included in the sign language dictionary information may include words of the following types: proper nouns, neologisms, slang, foreign words, and technical terms. Words corresponding to the predetermined types may include words that are included in the sign language dictionary information but whose clear meaning is difficult to convey. A method for determining the first word set and the second word set is described in detail with reference to FIGS. 5 and 6 .

[0091] Electronic devices can effectively convey meaning by providing sub-images along with sign language videos for words that must be expressed with fingers instead of sign language, or words that can be expressed with sign language but have difficulty conveying their clear meaning, thereby reducing the large amount of animation resources and time delay required when expressing the words with fingers. For example, if the original text contains 'Eiffel Tower', instead of sequentially presenting finger spelling animations corresponding to 'ㅇ', 'ㅔ', 'ㅍ', 'ㅔ', 'ㄹ', 'ㅌ', 'ㅏ', and 'ㅂ', the electronic device can generate sub-images representing 'Eiffel Tower'.

[0092] In operation 430, the electronic device may generate a sign language video based on one or more words included in the first word set.

[0093] According to one embodiment, the sign language video may include one or more sign language animations, each corresponding to one or more words included in the first word set.

[0094] An electronic device can generate a sign language video using pre-stored sign language dictionary information and pre-stored sign language animation information. A method for generating a sign language video is described in detail with reference to FIG. 7.

[0095] In operation 440, the electronic device may generate a sub-image of a sign language video based on one or more words included in a second word set. The sub-image may be generated based on at least one word among the one or more words included in the second word set.

[0096] According to one embodiment, the electronic device may generate one or more sub-images, each corresponding to one or more words included in the second word set.

[0097] According to one embodiment, the sub-image of the sign language video may include an integrated sub-image. The integrated sub-image may be generated based on two or more words among one or more words included in the second word set. The electronic device may generate an integrated sub-image corresponding to a plurality of words among one or more words included in the second word set.

[0098] An electronic device can generate sub-images of a sign language video using a generative AI model. A method for generating sub-images is described in detail with reference to FIGS. 9 and 10.

[0099] In one embodiment, operations 430 and 440 can be performed independently and in parallel.

[0100] In operation 450, the electronic device can generate a synthetic video including a sign language video and a sub-image.

[0101] An electronic device can generate a synthetic video by superimposing a sub-image onto a sign language video. A method for generating a synthetic video is described in detail with reference to FIG. 11.

[0102] The electronic device can output a composite video. The electronic device can display the composite video through a display (e.g., the display module (160) of FIG. 1).

[0103] FIG. 5 is a flowchart of a method for determining a first word set and a second word set according to one embodiment.

[0104] According to one embodiment, the operations 510 to 530 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1 or the processor (310) of FIG. 3) and a memory (e.g., the memory (130) of FIG. 1 or the memory (320) of FIG. 3).

[0105] According to one embodiment, operations 510 to 530 may be associated with operation 420 of determining the first word set and the second word set of FIG. 4. For example, operation 420 may include operations 510 to 530.

[0106] As described with reference to FIG. 4, the electronic device can determine analysis information for the original text. The analysis information for the original text may include at least one of the sentence types of the original text, the sentence components of words included in the original text, or the parts of speech of words included in the original text.

[0107] In operation 510, the electronic device can convert the original text into text suitable for the word order and grammar of the sign language based on the analysis information.

[0108] An electronic device can process the original text according to predetermined conditions based on the analysis information. The predetermined conditions may represent preset conditions (or rules) for converting the original text into text conforming to the word order and grammar of sign language. For example, the predetermined conditions may include converting or omitting at least a portion of the original text.

[0109] Electronic devices can convert the conjugated forms of words, phrases, or sentences in the original text into their base forms. If necessary, the device can convert passive to active voice, change word order, omit auxiliary verbs, or omit subjects as long as the meaning of the original text remains intact. The device can also omit words (e.g., endings) that, in context, can still convey the meaning.

[0110] Depending on the language of the source text, there may be differences in how the original text is converted. For example, if the original text is in Korean, an electronic device may remove particles and convert phrases containing conjugated endings to their base form. If the original text is in English, an electronic device may remove prepositions and omit contextually obvious auxiliary verbs.

[0111] In operation 520, the electronic device may determine, based on pre-stored sign language dictionary information, words included in the converted text that are not included in the sign language dictionary information or words that correspond to a predetermined type as a second word set.

[0112] Words not included in the sign language dictionary may include types of words such as proper nouns, neologisms, slang, foreign words, and technical terms. For example, an electronic device may determine words not included in the sign language dictionary, such as Eiffel Tower, Tokyo, smartwatch, selfie, YouTuber, drone, and macaron, as a second word set. The electronic device may also determine words that are not included in the sign language dictionary but can be expressed as other words or combinations of words included in the sign language dictionary, as a second word set. For example, words that belong to a subconcept (or subcategory) of at least one entry in the sign language dictionary, such as steak (beef + western food), air fryer (kitchen + product), cactus, tulip, and dandelion (plant), may be determined as a second word set. For words that belong to a subconcept of a entry, a sub-image expressing the specific meaning of the word may be more advantageous in conveying the meaning than a sign language gesture for the entry in a broader sense.

[0113] Words corresponding to predetermined types may include words that are included in sign language dictionary information but whose clear meaning is difficult to convey. According to one embodiment, words corresponding to predetermined types may be words associated with visible physical concepts among words included in sign language dictionary information. In particular, words corresponding to predetermined types may be words associated with physical concepts according to degree. For example, the electronic device may determine words associated with physical concepts such as color (e.g., red, dark red, reddish), size (e.g., big, small), shape (e.g., round, sharp), texture (e.g., smooth), contrast (e.g., bright, dark), indication (e.g., this, that, that), number (e.g., n pieces, nth), and quantity (e.g., many, few) as the second word set. For example, words associated with emotions, states, or abstract concepts that are difficult to visually confirm (e.g., scent, happiness, my, same, beautiful, dream, taste) may not be determined as the second word set.

[0114] In operation 530, the electronic device may determine a word included in the converted text that is not determined to be a second word set as a first word set for generating a sign language video.

[0115] According to one embodiment, the electronic device may correct a word modified by a word included in a second word set among one or more words included in a first word set to a second word set based on analysis information. 'Modification' may include not only one-to-one modification between words, but also a case where a word, phrase, or clause modifies another word, phrase, or clause. The electronic device may correct a word modified by a word included in a second word set (or a phrase or clause including a word belonging to the second word set) among one or more words included in the first word set to a second word set based on analysis information. Words associated with emotions, states, or abstract concepts that are difficult to visually confirm may not be corrected to the second word set.

[0116] According to one embodiment, the electronic device may correct a word that modifies a word in a second word set among one or more words included in a first word set to a second word set based on analysis information. 'Modification' may include not only one-to-one modification between words, but also a case where a word, phrase, or clause modifies another word, phrase, or clause. The electronic device may correct a word (or a word included in a phrase or clause) that modifies a word in a second word set (or a phrase or clause including a word belonging to the second word set) among one or more words included in the first word set to a second word set based on analysis information. Words associated with emotions, states, or abstract concepts that are difficult to visually confirm may not be corrected to the second word set.

[0117] For example, in a text containing 'apple' and '검빨언' modifying 'apple', 'apple' included in the sign language dictionary information can be determined as the first word set. '검빨언' can be determined as the second word set because it corresponds to a predetermined type associated with a visible physical concept. At this time, 'apple' can be corrected to the second word set because it is a word modified by a word included in the second word set. In a text containing 'apple' and '불그스레한' modifying 'apple', 'apple' can similarly be determined as the first word set, and '불그스레한' can be determined as the second word set. At this time, 'apple' can be corrected to the second word set. The electronic device can generate an integrated sub-image corresponding to words in a modifying relationship, such as 'apple' and '검빨언', and 'apple' and '불그스레한'. Therefore, the integrated sub-images for each of the "red apple" and "reddish apple" are visible and can effectively express different characteristics according to degree (or intensity). The method for generating the integrated sub-images is described in detail with reference to Fig. 10.

[0118] FIG. 6 is a diagram illustrating a method for determining a first word set and a second word set according to an example.

[0119] As described with reference to FIGS. 3 to 5, an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (300) of FIG. 3) can obtain original text for generating a sign language image. For example, the electronic device can obtain the original text, "Last night I saw fireworks at the Eiffel Tower and ate T-bone steak." The electronic device can determine analysis information about the original text.

[0120] Analysis information about the original text may include at least one of the sentence types of the original text, the sentence components of words included in the original text, or the parts of speech of words included in the original text.

[0121] The original text mentioned above is a compound sentence that connects the first main clause, "I saw fireworks at the Eiffel Tower last night," and the second main clause, "I ate a T-bone steak," with the coordinating conjunction "and." If the original text mentioned above is in Korean, "I" is the subject, "last" is an adverb, "night" is an adverb (or, "last night" is an adverbial phrase), "at the Eiffel Tower" is an adverb, "fireworks" is the object, "saw" is the predicate, "T-bone" is an adjective, "steak" is the object (or, "T-bone steak" is the object), and "ate" is the predicate. If the original text mentioned above is in Korean, the original text may include particles. On the other hand, if the original text is in English, the original text contains prepositions and articles, and 'I saw fireworks at the Eiffel Tower last night' and 'I ate steak' are predicates, and 'T-bone' is an adjective.

[0122] Based on the analysis information, the electronic device can convert the original text into text that conforms to the word order and grammar of sign language. For example, the electronic device can omit particles (or prepositions and articles) from the original text (e.g., from the Eiffel Tower to the Eiffel Tower), convert inflected expressions to their base form (e.g., ate to eat), and omit subjects.

[0123] The electronic device may determine, based on pre-stored sign language dictionary information, words included in the converted text that are not included in the sign language dictionary information or words that correspond to a predetermined type as a second word set. For example, the electronic device may determine, among the words included in the converted text that are not included in the sign language dictionary information, words such as "Eiffel Tower," "fireworks," "T-bone," and "steak," as a second word set. The electronic device may determine, among the words included in the converted text that are not included in the second word set, words such as "yesterday," "night," "see," and "eat," as a first word set for generating a sign language video.

[0124] Figure 7 is a flowchart of a method for generating sign language video according to one embodiment.

[0125] According to one embodiment, the operations 710 to 730 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1 or the processor (310) of FIG. 3) and a memory (e.g., the memory (130) of FIG. 1 or the memory (320) of FIG. 3).

[0126] In one embodiment, actions 710 to 730 may be associated with action 430 of generating a sign language video of FIG. 4. For example, action 430 may include actions 710 to 730.

[0127] In operation 710, the electronic device may determine one or more representative headings each corresponding to one or more words included in a first word set for generating a sign language video based on pre-stored sign language dictionary information.

[0128] Sign language dictionary information may include words that can be expressed in sign language. Sign language dictionary information may include relationships between representative headings and homonyms, and relationships between representative headings and compound words.

[0129] One or more words included in the first word set may be one of a representative headword mapped to a sign language animation, a headword (homonym) that can be expressed with the same sign language animation as the representative headword, or a compound word formed by a combination of representative headwords.

[0130] The electronic device may determine one or more representative headings corresponding to one or more words included in the first word set, based on the relationship between the representative headings and homonyms included in the sign language dictionary information, and the relationship between the representative headings and compound words. For example, if the first word included in the first word set is a heading that is not a representative heading, the electronic device may determine the first representative heading associated with the heading as the representative heading corresponding to the first word.

[0131] In operation 720, the electronic device can obtain one or more sign language animations each corresponding to one or more representative headings based on pre-stored sign language animation information.

[0132] The sign language animation information may include sign language animations mapped (or corresponding) to a plurality of representative headings among words included in the sign language dictionary information. Based on the sign language animation information, the electronic device may obtain one or more sign language animations mapped to one or more representative headings, each corresponding to one or more words included in the first word set.

[0133] In operation 730, the electronic device can generate a sign language video using one or more sign language animations.

[0134] The electronic device can generate a sign language video by sequentially combining one or more sign language animations according to the word order of the converted text (i.e., the original text is converted to suit the word order and grammar of the sign language based on the analysis information).

[0135] The electronic device can generate a sign language video so that the sign language movements of the virtual avatar corresponding to the end point of the first sign language animation and the start point of the second sign language animation that follows the first sign language animation are naturally connected. For example, the electronic device can blend the graphics of the virtual avatar corresponding to the end point of the first sign language animation and the start point of the second sign language animation. For example, the electronic device can generate an intermediate motion animation of the virtual avatar between the first and second sign language animations, and insert the intermediate motion animation between the first and second sign language animations.

[0136] FIG. 8 is a diagram illustrating a method for generating a sign language video according to one embodiment.

[0137] As described with reference to FIGS. 3 to 7, an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (300) of FIG. 3) can obtain original text for generating a sign language image. The electronic device can determine analysis information about the original text. Based on the analysis information, the electronic device can convert the original text into text suitable for the word order and grammar of sign language. Based on pre-stored sign language dictionary information, the electronic device can determine, among words included in the converted text, words that are not included in the sign language dictionary information or words that correspond to a predetermined type, as a second word set for generating a sub-image of a sign language video. The electronic device can determine, among words included in the converted text, words that are not determined as the second word set, as a first word set for generating a sign language video. For example, as described with reference to FIG. 6, the electronic device may determine 'Eiffel Tower', 'fireworks', 'T-bone', and 'steak' as the second word set, and 'yesterday', 'night', 'see', and 'eat' as the first word set.

[0138] The electronic device can determine one or more representative headings corresponding to one or more words included in a first word set for generating a sign language video based on pre-stored sign language dictionary information. Referring to FIG. 8, the one or more representative headings corresponding to one or more words included in the first word set are "yesterday," "night," "see," and "eat."

[0139] An electronic device can obtain one or more sign language animations corresponding to one or more representative headings based on pre-stored sign language animation information. Referring to FIG. 8, the electronic device can obtain sign language animations corresponding to "yesterday," "night," "see," and "eat," respectively.

[0140] An electronic device can generate a sign language video by sequentially combining one or more sign language animations according to the word order of the converted text (i.e., the original text has been converted to conform to the word order and grammar of sign language based on analysis information). Referring to FIG. 8, the electronic device can combine sign language animations in the order of "yesterday," "night," "see," and "eat."

[0141] FIG. 9 is a flowchart of a method for generating a sub-image according to one embodiment.

[0142] According to one embodiment, the operations 910 and 920 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1 or the processor (310) of FIG. 3) and a memory (e.g., the memory (130) of FIG. 1 or the memory (320) of FIG. 3).

[0143] In one embodiment, operations 910 and 920 may be associated with operation 440 of generating a sub-image of FIG. 4. For example, operation 440 may include operations 910 and 920.

[0144] According to one embodiment, the electronic device may generate one or more sub-images, each corresponding to one or more words included in the second word set.

[0145] In operation 910, the electronic device may determine a prompt using one or more words included in a second word set to generate a sub-image of the sign language video.

[0146] According to one embodiment, the electronic device can determine the prompt by inserting at least a portion of one or more words included in the second word set into a predetermined format. The predetermined format may be a format of a prompt requesting an image for each word included in the second word set, such as, for example, [Generate an image for each of 'one or more words included in the second word set'], [Generate an image representing 'the first word included in the second word set'], [Generate an image of 'the first word included in the second word set' and create an image of 'the second word included in the second word set'].

[0147] In operation 920, the electronic device may generate a sub-image using a generative AI model (e.g., the generative model (250) of FIG. 2) based on the prompt.

[0148] In one embodiment, the electronic device can generate a sub-image corresponding to the prompt using an on-device generated AI model of the electronic device.

[0149] In one embodiment, the electronic device may transmit a prompt to an external server (e.g., server (108) of FIG. 1) containing a generative AI model. The electronic device may obtain a sub-image corresponding to the prompt from the external server.

[0150] FIG. 10 is a flowchart of a method for generating a sub-image according to one embodiment.

[0151] According to one embodiment, the operations 1010 and 1020 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1 or the processor (310) of FIG. 3) and a memory (e.g., the memory (130) of FIG. 1 or the memory (320) of FIG. 3).

[0152] In one embodiment, operations 1010 and 1020 may be associated with operation 440 of generating a sub-image of FIG. 4. For example, operation 440 may include operations 1010 and 1020.

[0153] As described above with reference to FIG. 4, the sub-image may be generated based on at least one word among one or more words included in the second word set.

[0154] According to one embodiment, the sub-image of the sign language video may include an integrated sub-image. The integrated sub-image may be generated based on two or more words among one or more words included in the second word set. The electronic device may generate an integrated sub-image corresponding to a plurality of words among one or more words included in the second word set.

[0155] In operation 1010, the electronic device may determine a subset of words that satisfy an association condition among one or more words included in a second word set for generating a sub-image of a sign language video based on the analysis information.

[0156] In one embodiment, the relevance condition may include a condition where any of one or more words included in the second word set correspond to an object and an adverbial within a main clause or a subordinate clause, respectively. The object may include not only a single word but also an object phrase composed of two or more words. The adverbial may include not only a single word but also an adverbial phrase composed of two or more words.

[0157] In one embodiment, the association condition may include a condition in which any of one or more words included in the second word set are in a modifier relationship. The modifier relationship may include not only one-to-one modifiers between words, but also cases in which a word, phrase, or clause modifies another word, phrase, or clause.

[0158] At operation 1020, the electronic device can generate a unified sub-image corresponding to a subset of words.

[0159] In one embodiment, the electronic device may determine a prompt using a subset of words.

[0160] According to one embodiment, the electronic device can determine a prompt by inserting a subset of words from among one or more words included in the second word set into a predetermined format. The predetermined format may be a format of a prompt requesting an integrated image for a subset of words, such as, for example, [Generate an image in which 'a first word of the subset of words' and 'a second word of the subset of words' are represented], [Generate an image of 'a second word of the subset of words' in which features of 'a first word of the subset of words' are represented], [Generate a single image in which 'a first word of the subset of words' and 'a second word of the subset of words' appear together].

[0161] In one embodiment, the electronic device can generate an integrated sub-image using a generative AI model based on a prompt.

[0162] In one embodiment, the electronic device can generate an integrated sub-image corresponding to a prompt using an on-device generated AI model of the electronic device.

[0163] In one embodiment, the electronic device may transmit a prompt to an external server (e.g., server (108) of FIG. 1) that includes a generative AI model (e.g., generative model (250) of FIG. 2). The electronic device may obtain an integrated sub-image corresponding to the prompt from the external server.

[0164] FIG. 11 and FIG. 12 are drawings each illustrating a method for generating a sub-image according to an example.

[0165] As described with reference to FIGS. 3 to 10, an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (300) of FIG. 3) can obtain original text for generating a sign language image. The electronic device can determine analysis information about the original text. Based on the analysis information, the electronic device can convert the original text into text suitable for the word order and grammar of sign language. Based on pre-stored sign language dictionary information, the electronic device can determine, as a second word set for generating a sub-image of a sign language video, words included in the converted text that are not included in the sign language dictionary information or that correspond to a predetermined type. The electronic device can determine, as a first word set for generating a sign language video, words included in the converted text that are not determined as the second word set.

[0166] Referring to FIG. 11, for example, as described with reference to FIG. 6, the electronic device may determine 'Eiffel Tower', 'fireworks', 'T-bone', and 'steak' as a second word set, and 'yesterday', 'night', 'see', and 'eat' as a first word set.

[0167] According to one embodiment, the electronic device may generate one or more sub-images, each corresponding to one or more words included in the second word set.

[0168] The electronic device may determine a prompt by inserting at least a portion of one or more words included in the second word set into a predetermined format. For example, the electronic device may determine (or generate) a prompt requesting an image for each word included in the second word set, such as [Create an image of each of 'Eiffel Tower', 'fireworks', 'T-bone', and 'steak'].

[0169] The electronic device can generate sub-images based on the prompt using a generative AI model (e.g., the generative model (250) of FIG. 2). Referring to FIG. 11, sub-images corresponding to 'Eiffel Tower', 'fireworks', 'T-bone', and 'steak' included in the second word set can be generated, respectively.

[0170] In one embodiment, the electronic device can generate an integrated sub-image corresponding to a plurality of words among one or more words included in the second word set.

[0171] The electronic device can determine a subset of words that satisfy an association condition among one or more words included in the second word set based on the analysis information.

[0172] The relevance condition may include a condition in which any one or more words included in the second word set correspond to an object and an adverbial in one main clause or one subordinate clause, respectively. For example, the electronic device may determine "fireworks" and "Eiffel Tower," which correspond to an object and an adverbial in one main clause among "Eiffel Tower," "fireworks," "T-bone," and "steak," respectively, included in the second word set, as the first subset.

[0173] The association condition may include a condition in which any of the words in the second word set are in a modifier relationship. For example, the electronic device may determine that among the words "Eiffel Tower," "fireworks," "T-bone," and "steak" in the second word set, "T-bone" and "steak" are in a modifier relationship, and thus constitute the second subset.

[0174] An electronic device can generate an integrated sub-image corresponding to a subset of words. The electronic device can determine a prompt using the subset of words. That is, the electronic device can generate integrated sub-images corresponding to the first subset and the second subset, respectively. The electronic device can determine the prompt by inserting the first subset and the second subset into a predetermined format. For example, the electronic device can determine (or generate) a prompt that requests an integrated image for words included in the first subset and the second subset, respectively, such as [Generate an image representing 'Eiffel Tower' and 'fireworks'] and [Generate an image of 'steak' that exhibits the characteristics of 'T-bone'].

[0175] An electronic device can generate integrated sub-images using a generative AI model based on a prompt. Referring to FIG. 11, integrated sub-images corresponding to the first sub-set and the second sub-set can be generated, respectively.

[0176] Referring to FIG. 12, for example, an electronic device can obtain the original text, "I bought a diffuser that smells like flower." The electronic device can determine analysis information about the original text. If the original text is in Korean, "I" is the subject, "that smells like flower" is the noun clause, "diffuser" is the object, and "bought" is the predicate. If the original text is in Korean, the original text may include particles. If the original text is in English, the original text includes a preposition and an article, "I bought a diffuser that smells like flower" is the predicate, and "flower-scented" is the adjective clause.

[0177] Based on the analysis information, the electronic device can convert the original text into text that conforms to the word order and grammar of sign language. For example, the electronic device can omit particles (or prepositions and articles) from the original text, convert inflected expressions to their base form, omit subjects, and omit words that, in context, can still convey the meaning.

[0178] The electronic device may determine, based on pre-stored sign language dictionary information, words included in the converted text that are not included in the sign language dictionary or that correspond to a predetermined type as a second word set. For example, the electronic device may determine the word "diffuser," which is not included in the sign language dictionary information, as a second word set. The electronic device may determine the words "flower," "fragrance," and "buy," which are not included in the second word set, as a first word set.

[0179] Based on the analysis information, the electronic device can correct a word (or a word included in a phrase or clause) that modifies a word included in a second word set (or a phrase or clause that includes a word belonging to the second word set) among one or more words included in the first word set to the second word set. However, words associated with emotions, states, or abstract concepts that are difficult to visually confirm are not corrected to the second word set. Referring to FIG. 12, the electronic device can correct 'flower' included in a clause modifying 'diffuser' included in the second word set among one or more words included in the first word set to the second word set. 'Fragrance', like 'flower', is a word included in a clause modifying 'diffuser', but is not corrected to the second word set because it is a word associated with an abstract concept. Consequently, the electronic device can determine 'fragrance' and 'buy' as the first word set, and 'flower' and 'diffuser' as the second word set.

[0180] According to one embodiment, the electronic device may generate one or more sub-images, each corresponding to one or more words included in the second word set. Referring to FIG. 12, sub-images corresponding to 'flower' and 'diffuser' included in the second word set may be generated.

[0181] According to one embodiment, the electronic device may generate an integrated sub-image corresponding to a plurality of words among one or more words included in a second word set. Based on the analysis information, the electronic device may determine a subset of words that satisfy an association condition among one or more words included in the second word set. For example, the electronic device may determine the words "flower" and "diffuser," which are in a formula relationship among the words included in the second word set, as the subset of words.

[0182] An electronic device can generate an integrated sub-image corresponding to a subset of words. Referring to FIG. 12, an integrated sub-image corresponding to a subset of words including 'flower' and 'diffuser' can be generated.

[0183] FIG. 13 is a flowchart of a method for generating a synthetic video according to one embodiment.

[0184] According to one embodiment, the operations 1310 and 1320 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1 or the processor (310) of FIG. 3) and a memory (e.g., the memory (130) of FIG. 1 or the memory (320) of FIG. 3).

[0185] As described with reference to FIG. 8, the electronic device may determine one or more representative headings, each corresponding to one or more words included in a first word set for generating a sign language video, based on pre-stored sign language dictionary information. The electronic device may obtain one or more sign language animations, each corresponding to one or more representative headings, based on pre-stored sign language animation information. The electronic device may generate a sign language video using the one or more sign language animations.

[0186] In one embodiment, operations 1310 and 1320 may be associated with operation 450 of generating the composite video of FIG. 4. For example, operation 450 may include operations 1310 and 1320.

[0187] In operation 1310, the electronic device may determine a target sign language animation corresponding to a sub-image (or, integrated sub-image) among one or more sign language animations included in the sign language video.

[0188] According to one embodiment, if a sub-image is generated based on a word corresponding to the object among words included in the second word set, the electronic device may determine a sign language animation for a predicate associated with the object as the target sign language animation. The predicate associated with the object may be a predicate included in the same main clause or subordinate clause as the object.

[0189] According to one embodiment, if a sub-image is generated based on a word corresponding to an adverb among words included in a second word set, the electronic device may determine a sign language animation for a predicate associated with the adverb as the target sign language animation. The predicate associated with the adverb may be a predicate included in the same main clause or subordinate clause as the adverb.

[0190] According to one embodiment, the electronic device may determine, as the target sign language animation, a sign language animation for a subject, object, or complement modified by the adjective, if the sub-image is generated based on a word corresponding to an adjective among words included in the second word set.

[0191] In one embodiment, the electronic device may determine, as the target sign language animation, a sign language animation for a predicate contained within the same main clause or subordinate clause as the subset of words, when the integrated sub-image is generated based on a subset of words corresponding to an object and an adverbial within a single main clause or subordinate clause.

[0192] In one embodiment, the electronic device may determine a sign language animation for a predicate associated with the object as the target sign language animation if the integrated sub-image is generated based on a subset of words that include a word corresponding to the object.

[0193] In one embodiment, the electronic device may determine a sign language animation for a predicate associated with an adverb as the target sign language animation if the integrated sub-image is generated based on a subset of words that include a word corresponding to the adverb.

[0194] In one embodiment, if the original text is in a language other than Korean, such as English, the aforementioned 'predicate' may be understood as a 'verb', and the aforementioned 'determinant' may be understood as a 'modifier' or 'qualifier'.

[0195] In operation 1320, the electronic device can generate a synthetic video by superimposing a sub-image onto the sign language video from a start time to an end time of the target sign language animation.

[0196] According to one embodiment, when the target sign language animation corresponding to a plurality of sub-images is the same, the electronic device may generate a synthetic video by sequentially superimposing the plurality of sub-images on the sign language video from the start time to the end time of the target sign language animation. For example, when a plurality of sub-images are generated based on each of a plurality of objects or a plurality of adverbials included in a single main clause or subordinate clause, a sign language animation for a predicate associated with the plurality of objects or the plurality of adverbials (i.e., a predicate included in a single clause) may be determined as a common target sign language animation.

[0197] According to one embodiment, the electronic device can generate a synthetic video by superimposing the plurality of sub-images onto the sign language video side by side (or simultaneously) from the start time to the end time of the target sign language animation, when the target sign language animation corresponding to the plurality of sub-images is the same.

[0198] According to one embodiment, the electronic device can determine the display location of a sub-image corresponding to a target sign language animation based on pre-stored sign language animation information.

[0199] Sign language animation information may include sign language animations mapped to (or corresponding to) a plurality of representative headings among words included in the sign language dictionary information.

[0200] Sign language animation information may include information regarding the display positions of pre-designated sub-images corresponding to each sign language animation. The display positions of the sub-images may be designated as the area closest to the area where the sign language movements of the virtual avatar are expressed in the sign language animation.

[0201] In one embodiment, the electronic device can generate a synthetic video by superimposing a sub-image onto a determined display location of a sign language video.

[0202] FIG. 14 and FIG. 15 are drawings each illustrating a method for generating a synthetic video according to an example.

[0203] As described with reference to FIGS. 3 to 13, an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (300) of FIG. 3) can generate a sign language video based on one or more words included in a first word set and generate sub-images of the sign language video based on one or more words included in a second word set. For example, as described with reference to FIG. 8, the electronic device can generate a sign language video by sequentially combining sign language animations corresponding to 'yesterday', 'night', 'see', and 'eat' included in the first word set, respectively. As described with reference to FIG. 11, the electronic device can generate integrated sub-images corresponding to the first subset (e.g., 'fireworks' and 'Eiffel Tower') and the second subset (e.g., 'T-bone' and 'steak'), respectively.

[0204] According to one embodiment, the electronic device can determine a target sign language animation corresponding to a sub-image (or, integrated sub-image) among one or more sign language animations included in a sign language video.

[0205] If the integrated sub-image is generated based on a subset of words corresponding to objects and adverbs within a main clause or a subordinate clause, the electronic device may determine a sign language animation for a predicate included in the same main clause or subordinate clause as the subset of words as the target sign language animation. Referring to FIG. 14, the electronic device may determine the sign language animation for "see" as the target sign language animation corresponding to the integrated sub-image generated based on the first subset (e.g., "fireworks" and "Eiffel Tower").

[0206] If the integrated sub-image is generated based on a subset of words that include a word corresponding to the object, the electronic device may determine the sign language animation for the predicate associated with the object as the target sign language animation. Referring to FIG. 14, the electronic device may determine the sign language animation for "eat" as the target sign language animation corresponding to the integrated sub-image generated based on a second subset (e.g., "T-bone" and "steak").

[0207] An electronic device can determine the display position of a sub-image corresponding to a target sign language animation based on pre-stored sign language animation information. The sign language animation information can include information on the display position of a pre-designated sub-image corresponding to each sign language animation. The display position of the sub-image can be designated as an area closest to the area where the sign language movement of the virtual avatar is expressed in the sign language animation. Referring to FIG. 15, the display positions (1501, 1502, 1503) of the sub-images corresponding to the sign language animations (1510, 1520, 1530) are illustrated, respectively.

[0208] The electronic device can generate a synthetic video by superimposing sub-images on a sign language video from a start time to an end time of a target sign language animation. Referring to FIG. 14, the electronic device can generate a synthetic video by superimposing integrated sub-images generated based on a first subset (e.g., 'fireworks' and 'Eiffel Tower') on a sign language video from a start time to an end time of a sign language animation for 'see', and superimposing integrated sub-images generated based on a second subset (e.g., 'T-bone' and 'steak') on a sign language video from a start time to an end time of a sign language animation for 'eat'.

[0209] In one embodiment, if the original text is in a language other than Korean, such as English, the aforementioned 'predicate' may be understood as a 'verb', and the aforementioned 'determinant' may be understood as a 'modifier' or 'qualifier'.

[0210] In one embodiment, a method for providing a sign language video, performed by an electronic device (101; 300), may include: an operation of determining analysis information for an original text; an operation of determining a first word set for generating a sign language video and a second word set for generating a sub-image of the sign language video from among words included in the original text based on the analysis information; an operation of generating a sign language video based on one or more words included in the first word set; an operation of generating a sub-image of the sign language video based on one or more words included in the second word set; and an operation of generating a synthetic video including the sign language video and the sub-image.

[0211] According to one embodiment, the analysis information may include at least one of a sentence type of the original text, a sentence component of words included in the original text, or a part of speech of words included in the original text.

[0212] According to one embodiment, the operation of determining a first word set and a second word set among words included in an original text based on analysis information may include: an operation of converting the original text into text suitable for word order and grammar of sign language based on the analysis information; an operation of determining, based on pre-stored sign language dictionary information, words included in the converted text that are not included in sign language dictionary information or that correspond to a predetermined type as a second word set for generating a sub-image of a sign language video; and an operation of determining, among words included in the converted text, words that are not determined as the second word set as the first word set for generating a sign language video.

[0213] According to one embodiment, the operation of determining a first word set and a second word set among words included in the original text based on the analysis information may further include an operation of correcting a word modified by a word included in the second word set among one or more words included in the first word set to the second word set based on the analysis information.

[0214] According to one embodiment, the operation of determining a first word set and a second word set among words included in the original text based on the analysis information may further include an operation of correcting a word that modifies a word included in the second word set among one or more words included in the first word set to the second word set based on the analysis information.

[0215] According to one embodiment, the operation of generating a sign language video based on one or more words included in a first word set may include the operation of determining one or more representative headings corresponding to one or more words included in the first word set based on pre-stored sign language dictionary information; the operation of obtaining one or more sign language animations corresponding to one or more representative headings based on pre-stored sign language animation information; and the operation of generating a sign language video using the one or more sign language animations.

[0216] In one embodiment, the operation of generating a sub-image of a sign language video based on one or more words included in the second word set may include: determining a prompt using one or more words included in the second word set; and generating a sub-image based on the prompt using a generative artificial intelligence (AI) model.

[0217] According to one embodiment, the operation of generating a sub-image of a sign language video based on one or more words included in a second word set may include the operation of determining a subset of words that satisfy an association condition among one or more words included in the second word set based on analysis information; and the operation of generating an integrated sub-image corresponding to the subset of words.

[0218] According to one embodiment, the relevance condition for determining a subset of words among one or more words included in the second word set may include a condition that any word among one or more words included in the second word set corresponds to an object and an adverbial in one independent clause or one dependent clause, respectively.

[0219] According to one embodiment, the association condition for determining a subset of words among one or more words included in the second word set may include a condition that any words among one or more words included in the second word set are in a modifier relationship.

[0220] In one embodiment, the operation of generating an integrated sub-image corresponding to a subset of words that satisfy an association condition among one or more words included in a second word set may include the operation of determining a prompt using the subset of words; and the operation of generating an integrated sub-image using a generative AI model based on the prompt.

[0221] According to one embodiment, the operation of generating a synthetic video including a sign language video and a sub-image may include the operation of determining a target sign language animation corresponding to a sub-image among one or more sign language animations included in the sign language video; and the operation of generating the synthetic video by superimposing the sub-image onto the sign language video from a start time to an end time of the target sign language animation.

[0222] According to one embodiment, the operation of generating a synthetic video by superimposing a sub-image on a sign language video may include the operation of determining a display position of a sub-image corresponding to a target sign language animation based on pre-stored sign language animation information; and the operation of generating the synthetic video by superimposing the sub-image on the determined display position of the sign language video.

[0223] According to one embodiment, a computer program may be stored in a computer-readable recording medium to execute a method for providing sign language video in combination with hardware.

[0224] According to one embodiment, a non-transitory computer-readable recording medium may store one or more programs including instructions, which, when individually or collectively executed by at least one processor (120; 310) of an electronic device (101; 300), cause the electronic device (101; 300) to: determine analysis information for an original text; determine a first word set for generating a sign language video from among words included in the original text based on the analysis information and a second word set for generating a sub-image of the sign language video; generate the sign language video based on one or more words included in the first word set; generate a sub-image of the sign language video based on one or more words included in the second word set; generate a composite video including the sign language video and the sub-image; and display the composite video through a display (160) of the electronic device (101; 300).

[0225] In one embodiment, an electronic device (101; 300) comprises at least one processor (120; 310) including processing circuitry; and a memory (130; 320) including one or more storage media for storing instructions, wherein when the instructions are individually or collectively executed by the at least one processor (120; 310), the electronic device (101; 300) may: determine analysis information for an original text; determine a first word set for generating a sign language video from among words included in the original text based on the analysis information and a second word set for generating a sub-image of the sign language video; generate the sign language video based on one or more words included in the first word set; generate a sub-image of the sign language video based on one or more words included in the second word set; and generate a synthetic video including the sign language video and the sub-image.

[0226] According to one embodiment, the analysis information may include at least one of a sentence type of the original text, a sentence component of words included in the original text, or a part of speech of words included in the original text.

[0227] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120; 310), the electronic device (101; 300) may: convert an original text into text suitable for word order and grammar of sign language based on analysis information, and, based on pre-stored sign language dictionary information, determine words included in the converted text that are not included in the sign language dictionary information or that correspond to a predetermined type as a second word set for generating a sub-image of a sign language video, and determine words included in the converted text that are not determined as the second word set as a first word set for generating a sign language video.

[0228] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120; 310), the electronic device (101; 300) may: correct a word modified by a word included in a second word set among one or more words included in a first word set to the second word set based on the analysis information.

[0229] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120; 310), the electronic device (101; 300) may: correct a word that modifies a word included in a second word set among one or more words included in a first word set to the second word set based on the analysis information.

[0230] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120; 310), the electronic device (101; 300) may be configured to: determine a prompt using one or more words included in a second word set; and generate a sub-image using a generative artificial intelligence (AI) model based on the prompt.

[0231] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing unit may execute an operating system (OS) and software applications running on the operating system. Furthermore, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0232] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.

[0233] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.

[0234] The hardware device described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.

[0235] Although the embodiments have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the described embodiments. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0236] Therefore, other implementations, embodiments and equivalents to the claims also fall within the scope of the claims described below.

Claims

1. In a method performed by an electronic device (101; 300), The action of determining analytical information about the original text; An operation of determining a first word set for generating a sign language video and a second word set for generating a sub-image of the sign language video from among words included in the original text based on the above analysis information; An operation of generating a sign language video based on one or more words included in the first word set; An operation of generating a sub-image of the sign language video based on one or more words included in the second word set; An operation of generating a synthetic video including the sign language video and the sub-image; and An operation of displaying the synthetic video through the display (160) of the electronic device (101; 300) including, method.

2. In paragraph 1, The above analysis information is, Including at least one of the sentence types of the original text, the sentence components of the words included in the original text, or the parts of speech of the words included in the original text. method.

3. In paragraph 1 or 2, The operation of determining the first word set and the second word set among the words included in the original text based on the analysis information is as follows: An operation of converting the original text into text suitable for the word order and grammar of sign language based on the above analysis information; An operation of determining, based on pre-stored sign language dictionary information, words included in the converted text that are not included in the sign language dictionary information or that correspond to a predetermined type as the second word set for generating the sub-image of the sign language video; and An operation of determining a word included in the converted text that is not determined as the second word set as the first word set for generating the sign language video. including, method.

4. In any one of paragraphs 1 to 3, The operation of determining the first word set and the second word set among the words included in the original text based on the analysis information is as follows: An operation of correcting a word modified by a word included in the second word set among one or more words included in the first word set based on the analysis information to the second word set. including more, method.

5. In any one of paragraphs 1 to 4, The operation of determining the first word set and the second word set among the words included in the original text based on the analysis information is as follows: An operation of correcting a word that modifies a word included in the second word set among one or more words included in the first word set based on the analysis information to the second word set. including more, method.

6. In any one of paragraphs 1 to 5, The operation of generating the sign language video based on one or more words included in the first word set is: An operation of determining one or more representative headings corresponding to one or more words included in the first word set based on pre-stored sign language dictionary information; An operation of obtaining one or more sign language animations corresponding to each of the one or more representative headings based on pre-stored sign language animation information; and An action of generating the sign language video using one or more of the above sign language animations. including, method.

7. In any one of paragraphs 1 to 6, The operation of generating the sub-image of the sign language video based on one or more words included in the second word set is: An action of determining a prompt using one or more words included in the second word set; and An action to generate the sub-image using a generative AI (artificial intelligence) model based on the above prompt. including, method.

8. In any one of paragraphs 1 to 7, The operation of generating the sub-image of the sign language video based on one or more words included in the second word set is: An operation of determining a subset of words that satisfy an association condition among one or more words included in the second word set based on the analysis information; and An operation to generate an integrated sub-image corresponding to a subset of the above words. including, method.

9. In any one of paragraphs 1 to 8, The relevance condition for determining a subset of words among one or more words included in the second word set is: Any of the words included in the second word set include a condition corresponding to an object and an adverb in one independent clause or one dependent clause, method.

10. In any one of paragraphs 1 to 9, The relevance condition for determining a subset of words among one or more words included in the second word set is: A condition in which any of one or more words included in the second word set are in a formula relationship, method.

11. In any one of paragraphs 1 to 10, An operation of generating an integrated sub-image corresponding to a subset of words that satisfy an association condition among one or more words included in the second word set is as follows: An action to determine a prompt using a subset of the above words; and An operation of generating the integrated sub-image using a generative AI model based on the above prompt. including, method.

12. In any one of paragraphs 1 to 11, The operation of generating the synthetic video including the sign language video and the sub-image is as follows: An operation of determining a target sign language animation corresponding to the sub-image among one or more sign language animations included in the sign language video; and An operation of generating the synthetic video by superimposing the sub-image onto the sign language video from the start time to the end time of the target sign language animation. including, method.

13. In any one of paragraphs 1 to 12, The operation of generating the synthetic video by superimposing the sub-image on the sign language video is as follows: An operation of determining a display position of the sub-image corresponding to a target sign language animation based on pre-stored sign language animation information; and An operation of generating the synthetic video by superimposing the sub-image on the determined display position of the sign language video. including, method.

14. In a non-transitory computer-readable recording medium, Store one or more programs containing instructions, When the above instructions are individually or collectively executed by at least one processor (120; 310) of the electronic device (101; 300), the electronic device (101; 300) causes: Determine analytical information about the original text, Based on the above analysis information, a first word set for generating a sign language video and a second word set for generating a sub-image of the sign language video are determined from among the words included in the original text, Generate a sign language video based on one or more words included in the first word set, Generating a sub-image of the sign language video based on one or more words included in the second word set, Generating a synthetic video including the above sign language video and the above sub-image, Displaying the synthetic video through the display (160) of the electronic device (101; 300) To do, Non-transitory computer-readable recording medium.

15. In the electronic device (101; 300), At least one processor (120; 310) comprising processing circuitry; and A memory (130; 320) comprising one or more storage media for storing instructions, When the above instructions are individually or collectively executed by at least one processor (120; 310), the electronic device (101; 300) causes: Determine analytical information about the original text, Based on the above analysis information, a first word set for generating a sign language video and a second word set for generating a sub-image of the sign language video are determined from among the words included in the original text, Generate a sign language video based on one or more words included in the first word set, Generating a sub-image of the sign language video based on one or more words included in the second word set, Generating a synthetic video including the above sign language video and the above sub-image, Displaying the synthetic video through the display (160) of the electronic device (101; 300) To do, Electronic devices (101; 300).

Citation Information

Patent Citations

  • Sign language video-generating device, sign language video-outputting device, sign language video-generating method, and program

    JP2011123422A

  • Translation device and program

    JP2021196708A

  • Translation apparatus and program

    JP2022164367A

  • Apparatus for cleaning water pipe using water turbine

    KR1020250157850A

  • Automated generation and presentation of sign language avatars for video content

    US11935170B1