System

A system converts music lyrics into sign language and generates synchronized avatar animations to increase interest and accessibility of sign language, addressing the lack of engagement with sign language and musical experiences for the hearing-blind.

JP2026022542APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024124059
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

There are limited opportunities for the general public to learn about sign language and for the hearing-blind to enjoy music, with few accessible methods to engage with sign language.

Method used

A system that converts music lyrics into sign language, generates avatar movements based on the sign language, and synchronizes the movements with the music rhythm to create animations, which are then distributed via a network.

Benefits of technology

Provides opportunities for people to learn sign language and enhances the musical experience for the hearing-blind by making sign language more accessible and visible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022542000001_ABST
    Figure 2026022542000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: This system includes a means for converting the lyrics of music into sign language, a means for generating the action of an avatar on the basis of the sign language, a means for generating a moving image by matching the generated action of the avatar with the rhythm of the music, and a means for distributing the generated moving image through a network.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Although sign language is sometimes featured on television and in dramas, there are few opportunities for the general public to learn about or come into contact with it. As a result, there is a need for ways to raise interest in sign language. Furthermore, there are limited ways for the hearing-blind to enjoy music. The purpose of this invention is to solve these problems, make sign language more accessible and visible, and provide opportunities for the hearing-blind to enjoy music. [Means for solving the problem]

[0005] The present invention provides a system that converts music lyrics into sign language, generates avatar movements based on the sign language, and generates animations that synchronize the generated avatar movements with the rhythm of the music. This system is characterized by including: means for inputting text data of lyrics; means for analyzing the text data of lyrics and converting it into sign language gesture data; means for executing a script that analyzes the sign language gesture data and generates an avatar movement sequence; means for generating animations that synchronize the generated avatar movements with the rhythm of the music; and means for distributing the generated animations via a network.

[0006] This will provide people who want to learn sign language and those with hearing and visual impairments with the opportunity to learn sign language through music, thereby raising interest in sign language.In addition, because the content is distributed in a format that is easy for users to view, it will increase opportunities for people who are unfamiliar with sign language to come into contact with sign language.

[0007] "Lyrics" are words or sentences used in a musical composition.

[0008] "Sign language" is a means of communication using visual gestures and movements used by people who are deaf or hard of hearing.

[0009] An "avatar" is a virtual character or persona created by a user or system in a digital environment.

[0010] An "action sequence" defines the temporal order or pattern in which an avatar performs a particular gesture or movement.

[0011] An "AI model" is a mathematical model or algorithm that has been trained to perform a specific task using artificial intelligence.

[0012] "Gesture data" refers to digital information that represents specific actions or poses in sign language.

[0013] A "network" is a mechanism for sending and receiving data using a communication system such as the Internet.

[0014] "Video" refers to digital content that expresses movement by displaying a series of still images at a constant speed.

[0015] A "system" is an entire configuration in which multiple elements work together to achieve a specific function or service. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing in accompaniment with the sign language. This system performs a series of processes, including inputting lyrics text, converting them into sign language, generating avatar movements, generating videos, and distributing them.

[0038] First, the user inputs the lyrics of the music into the system using their own device. For example, the user can copy the lyrics from a music distribution service app and paste them into the system's dedicated input form. Then, the process begins by clicking the "Sign Language Conversion" button.

[0039] Next, the device sends the lyrics text data to the server, where it converts the input lyrics data into JSON format and sends it to the server using a dedicated communication protocol for fast and reliable data transfer.

[0040] The server analyzes the received lyrics text and inputs it into a pre-trained AI model, which then analyzes the lyrics text and generates corresponding sign language gesture data, which is then used to generate avatar movements.

[0041] The server generates an avatar movement sequence based on this sign language gesture data, which specifically instructs the movement of each joint and body part of the avatar.The server then analyzes the music file to synchronize the avatar movement sequence with the rhythm and tone of the music, and combines a timeline with the rhythm.

[0042] The server then generates a dance video of the avatar based on the movement sequence, matching the avatar's movements with the music to create a visually appealing video. The video is saved in a common video format, such as MP4.

[0043] The server then distributes the generated video over the network, uploading the video file to a dedicated social networking platform, such as a video content sharing application.

[0044] Finally, users can watch the video on social media platforms. The video is presented in a short format, making it easy to watch and share. The video also includes relevant metadata (e.g., original lyrics, sign language meanings, etc.), making it easy for viewers to understand the sign language content.

[0045] As a concrete example, if a user inputs the lyrics "Come Spring" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on that, generates a dance video of the avatar in time with the rhythm of the music, and finally distributes it to a social media platform.

[0046] The implementation of this system will enable people with visual and hearing impairments to learn sign language and to experience music, and is expected to increase interest in sign language.

[0047] The processing flow will be explained below.

[0048] Step 1:

[0049] The user inputs the lyrics of a song into the system. The user copies the lyrics from a music distribution service app and pastes them into the system's dedicated input form. Then, the user clicks the "Convert to Sign Language" button.

[0050] Step 2:

[0051] The device sends lyric text data to the server. The device converts the input lyric data into a JSON format request and sends it to the server.

[0052] Step 3:

[0053] The server receives the lyrics text and inputs it into the AI ​​model. The server analyzes the lyrics text and issues a request to the AI ​​model's API. The AI ​​model analyzes the lyrics and generates corresponding sign language gesture data.

[0054] Step 4:

[0055] The server receives the generated sign language data and stores it in an internal database, which is used to generate movement sequences for the avatar.

[0056] Step 5:

[0057] The server generates an avatar's movement sequence from the sign language data. The server executes a script that generates a sequence corresponding to the avatar's joint and body movements based on the sign language gesture data.

[0058] Step 6:

[0059] The server generates a dance video of the avatar based on the sign language sequence and synchronizes it with the rhythm and tone of the music. The server analyzes the music file and calculates the timing based on the rhythm and tone. The server then reflects the sign language sequence in the avatar's movements according to the timing and generates the animation using 3D modeling software.

[0060] Step 7:

[0061] The server saves the generated dance video data and prepares it for distribution. The server saves the generated dance video in a format such as MP4. The server converts the video file into a format suitable for distribution on social media and adds metadata (e.g., original lyrics, sign language explanations).

[0062] Step 8:

[0063] The server uploads the created dance video to the SNS platform. The server then sends the saved video file to the SNS platform's API for uploading.

[0064] Step 9:

[0065] Allow users to watch videos on the social media platform, which then makes the uploaded videos public and available for users to watch.

[0066] Step 10:

[0067] To make it easier for users to watch, the videos are distributed in short video format and related metadata (e.g., original lyrics, meaning of sign language) is also provided. The server automatically adds the video file metadata to the video description section of the social media platform.

[0068] Example 1

[0069] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0070] With conventional technology, it has been difficult to convert music lyrics into sign language and efficiently generate and distribute videos of avatars dancing using sign language. Furthermore, there have been limited ways for the hearing-blind to enjoy music, and there has also been a lack of methods to promote learning sign language. This has limited musical experiences and learning opportunities for them.

[0071] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0072] In this invention, the server includes means for converting lyric text data into JSON format and sending it to the server, means for analyzing the lyric text data using an AI model and converting it into sign language gesture data, and means for generating an avatar movement sequence based on the sign language gesture data. This allows the generation of an avatar movement sequence in sync with the rhythm of music, providing a means for the hearing-blind to enjoy music and promoting the learning of sign language.

[0073] "Lyric text data" refers to data in which music lyrics are saved as text information.

[0074] "Sign language gesture data" is data that expresses each gesture or movement in sign language.

[0075] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format that consists of data in key-value pairs.

[0076] An "AI model" is an algorithm or system trained using artificial intelligence techniques to perform a specific task or analysis.

[0077] An "avatar" refers to a virtual person or character in a digital environment.

[0078] A "movement sequence" is a series of commands that specify the movement of each joint and body part of the avatar along a time axis.

[0079] "3D modeling software" is software for creating, editing, and displaying three-dimensional models.

[0080] The "MP4 format" is a common digital file format for storing video and audio.

[0081] A "network" is a system of computers and other devices connected to exchange data with each other.

[0082] An "SNS platform" is an online service that allows users to post, share, and view content.

[0083] "Sign language" is a system of visual gestures and movements that enables deaf people to communicate.

[0084] The present invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing in accompaniment with the sign language. An embodiment of this system will be described in detail below.

[0085] First, the user inputs the lyrics of the music using their own device. The user copies the lyrics from a music distribution service application or similar and pastes them into the designated input form. Then, the process begins by clicking the "Sign Language Conversion" button.

[0086] Next, the device converts the entered lyrics text data into JSON format and sends it to the server. The device uses a dedicated communication protocol for this data conversion and transmission.

[0087] The server analyzes the received lyric text data using a pre-trained generative AI model, which converts the lyric text into sign language gesture data, which is then used to generate a movement sequence for the avatar.

[0088] The server then generates a motion sequence for the avatar based on the sign language gesture data, which specifically instructs the movements of each joint and body part of the avatar. It also analyzes the music file and combines a timeline with the rhythm of the music to synchronize the avatar's movements with the rhythm of the music.

[0089] The server then generates a dance video of the avatar based on the movement sequence, using 3D modeling software. The video, which matches the avatar's movements with the music, is visually appealing and saved in a common video format such as MP4.

[0090] Finally, the server distributes the generated video over the network. The video file is uploaded to a dedicated social networking platform, where users can watch the video. The video also includes relevant metadata, such as the original lyrics and sign language meanings, making it easy for viewers to understand the sign language content.

[0091] For example, if a user inputs the lyrics "Spring Comes" into the system and clicks the "Sign Language Conversion" button, the server converts the lyrics into sign language and generates sign language gesture data. Next, based on the generated sign language gesture data, the server creates a movement sequence for the avatar and generates a video of the avatar dancing to the rhythm of the music. Finally, this video is distributed to a social networking platform.

[0092] Example prompt sentence:

[0093] "Enter the following lyrics, 'Spring Comes,' into the system, convert them into sign language, generate a dance video with an avatar, and post it on social media."

[0094] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0095] Step 1:

[0096] A user uses his / her terminal to input the lyrics text of the music.

[0097] Specific operation: The user copies lyrics from a music distribution service application, pastes them into a dedicated input form, and clicks the "Convert to sign language" button.

[0098] Input: Music lyric text

[0099] Output: Trigger the start of processing by clicking the "Sign Language Conversion" button

[0100] Step 2:

[0101] The device converts the entered lyrics text data into JSON format and sends it to the server.

[0102] Specific operation: The terminal converts the text data into JSON format and sends it to the server using a dedicated communication protocol.

[0103] Input: Lyric text entered by the user

[0104] Output: Lyrics text data in JSON format sent to the server

[0105] Step 3:

[0106] The server analyzes the received lyrics text data.

[0107] Specific operation: The server analyzes the received JSON format lyrics text data using the analysis module.

[0108] Input: Lyrics text data in JSON format

[0109] Output: Parsed lyrics text data

[0110] Step 4:

[0111] The server inputs the analyzed lyrics text data into a generative AI model to generate sign language gesture data.

[0112] Specific operation: Text data is input into the generative AI model, and sign language gesture data is generated.

[0113] Input: Parsed lyrics text data

[0114] Output: Sign language gesture data

[0115] Step 5:

[0116] The server generates a movement sequence for the avatar based on the sign language gesture data.

[0117] Specific movements: Based on sign language gesture data, a movement sequence is generated that specifically instructs the movement of each joint and body part of the avatar.

[0118] Input: Sign language gesture data

[0119] Output: Avatar movement sequence

[0120] Step 6:

[0121] The server analyzes the music files and combines the movement sequences with rhythm.

[0122] Specific Actions: Analyze music files and combine action sequences into a timeline based on the rhythm of the music.

[0123] Input: Avatar movement sequence, music file

[0124] Output: Motion sequence synchronized with the music rhythm

[0125] Step 7:

[0126] The server generates a dance video based on the movement sequence.

[0127] Specific Movement: Using 3D modeling software, a dance video is generated that combines an avatar's movement sequence with music.

[0128] Input: A sequence of movements synchronized with the musical rhythm

[0129] Output: Generated dance video (MP4 format, etc.)

[0130] Step 8:

[0131] The server distributes the generated video over the network.

[0132] Specific operation: Upload the generated video file to a social media platform.

[0133] Input: Generated dance video (e.g. MP4 format)

[0134] Output: Upload completion notification to SNS platform

[0135] Step 9:

[0136] Users watch videos on social media platforms.

[0137] Specific Action: A user visits a social media platform to watch and share a video.

[0138] Input: Video link on social media platform

[0139] Output: Videos watched by users, videos shared

[0140] (Application example 1)

[0141] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0142] Conventional food delivery services have the problem that it is difficult for visually or hearing impaired users to understand the details of the dishes. In addition, there are many cases where explanations are not available in sign language, and a sign language interpreter is required. There is a need for a means to enable visually or audibly impaired users to understand detailed descriptions of dishes.

[0143] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0144] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for generating a video in which the generated avatar movements are synchronized with the rhythm of the music, means for distributing the generated video via a network, means for converting text describing dishes into sign language and generating a video in which an avatar uses sign language, and means for distributing the generated sign language video to visually impaired users, thereby enabling visually or hearing impaired users to understand the details of dishes when using a food delivery service.

[0145] "Musical lyrics" are text, usually consisting of words, that is sung to accompany a melody in a musical performance or song.

[0146] "Sign language" is a visual language that primarily communicates through hand and finger movements and facial expressions.

[0147] An "avatar" is a character or graphical representation that represents a user and acts on behalf of the user in a virtual space or on a computer.

[0148] "Generating behavior" refers to creating a series of actions or poses for an avatar or character on a computer.

[0149] "Musical rhythm" refers to the temporal pattern or tempo of music.

[0150] "Generating video" refers to creating a moving image by displaying a series of images or frames in succession.

[0151] "Distributing" refers to transmitting digital content over a network and making the content viewable on the receiving end.

[0152] "Description text of a dish" is a text about the contents, characteristics, ingredients, cooking method, etc. of a dish.

[0153] "Sign language gesture data" is a digital representation of sign language movements, and is typically data that indicates the position, direction, and movement of the hands and fingers.

[0154] "Visually impaired people" refers to people who have impaired vision and who have difficulty or are unable to understand visual information.

[0155] "Hearing impaired people" refers to people who have hearing impairments and who have difficulty or are unable to understand auditory information.

[0156] An "AI model" refers to an algorithm or computational model built using machine learning or deep learning, which performs data analysis and predictions.

[0157] "Via a network" refers to sending and receiving data via a communications infrastructure such as the Internet or a local area network.

[0158] This invention provides a system that enables visually and hearing impaired users to understand the details of dishes in food delivery services. Specifically, we apply technology that converts music lyrics into sign language to build a system that converts text describing dishes into sign language and generates and distributes videos using sign language by an avatar.

[0159] Basic system configuration

[0160] The server has the following functions:

[0161] 1. Text analysis:

[0162] The system analyzes the text description of a dish entered by the user and converts it into sign language gesture data using a generative AI model (using PyTorch and TensorFlow).

[0163] 2. Avatar movement generation:

[0164] Based on the analyzed sign language gesture data, a movement sequence for the avatar is generated, which specifically specifies the movements of each joint and body part required for sign language.

[0165] 3. Speech generation:

[0166] Use Google Cloud Speech-to-Text and Amazon Polly to convert the text of the cooking instructions into an audio file and add audio to the sign language video.

[0167] 4. Video Generation:

[0168] Using a video editing API such as OpenCV, sign language videos are generated so that the avatar's movements match the rhythm of the music.

[0169] 5. Delivery:

[0170] Using AWS S3 and CloudFront, the generated sign language videos are distributed to visually and hearing impaired people within a food delivery service application.

[0171] The user terminal has the following functions:

[0172] 1. Input interface:

[0173] It provides an interface for inputting text describing a dish. The user inputs the text and clicks the "Convert to sign language" button to start the process.

[0174] 2. Watching videos:

[0175] The distributed sign language video is played within the application, allowing the content to be conveyed to visually and hearing impaired people.

[0176] Processing flow

[0177] The user enters a text description of the dish: "This curry is made with special spices and is mild." The user's device sends the entered text data to the server, which converts the text into sign language gesture data. Next, the avatar movement generation function generates an avatar movement sequence based on the sign language gesture data. Furthermore, the audio file generated using Google Cloud Speech-to-Text and Amazon Polly is used to generate a sign language video of the avatar using a video editing API such as OpenCV. Finally, this video is distributed via AWS S3 and CloudFront and played on the user's device.

[0178] Prompt Sentence Examples

[0179] Below are some examples of dish descriptions:

[0180] "This curry is made with special spices and is not too spicy."

[0181] This pasta is made with plenty of seasonal vegetables.

[0182] This will enable users with visual or hearing impairments to understand details about dishes through food delivery services.

[0183] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0184] Step 1: User Input

[0185] The user inputs a text description of a dish through the food delivery service's application. For example, they input "This curry is made with special spices and is mild," and then click the "Sign Language Conversion" button. The input is sent to the device as text data.

[0186] Step 2: Send text data

[0187] The terminal converts the input text data into JSON format and sends it to the server. An appropriate communication protocol is used to ensure accurate data transfer. The input is text data and the output is JSON format data.

[0188] Step 3: Generate sign language gesture data

[0189] The server analyzes the received text data and inputs it into a generative AI model (using PyTorch or TensorFlow). The model converts the text into sign language gesture data. For example, it generates sign language gesture data for "curry." The input is text data, and the output is sign language gesture data.

[0190] Step 4: Generate avatar movement sequences

[0191] The server analyzes the sign language gesture data and generates a movement sequence for the avatar. A script is generated that specifically instructs the movement of each joint and body part, and the avatar's movement is determined based on this. The input is the sign language gesture data, and the output is the avatar's movement sequence.

[0192] Step 5: Speech generation

[0193] The server converts the text of the cooking instructions into an audio file using Google Cloud Speech-to-Text and Amazon Polly. The generated audio file is used to accompany the sign language video. The input is text data, and the output is an audio file.

[0194] Step 6: Video Generation

[0195] The server uses a video editing API such as OpenCV to combine the avatar's movement sequence based on the sign language gesture data with the generated audio file to generate a sign language video. This video is visually easy to understand, with the sign language and audio synchronized. The input is the avatar's movement sequence and audio file, and the output is the final sign language video.

[0196] Step 7: Publish your video

[0197] The server distributes the generated sign language video using AWS S3 and CloudFront, allowing users to watch the video on their devices. The input is the sign language video file, and the output is the video URL distributed over the network.

[0198] Step 8: Watch the video

[0199] The user device plays the sign language video delivered from the server and communicates the details of the dish to the visually and hearing impaired. The user can watch the video within the app and understand the explanation. The input is the video URL, and the output is the sign language video that is played.

[0200] This allows users with visual or hearing impairments to understand the details of the dishes and make appropriate choices.

[0201] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0202] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[0203] A specific embodiment of this system is described below.

[0204] First, the user inputs the lyrics of a song into the system. The user copies the lyrics from the music distribution service app and pastes them into the system's dedicated input form. The process begins by clicking the "Convert to Sign Language" button.

[0205] Next, the device sends the lyrics text data to the server, converts the input lyrics data into a JSON format request, and sends it to the server.

[0206] The server analyzes the received lyrics text and inputs it into the AI ​​model, which then analyzes the lyrics text and generates corresponding sign language gesture data, which is then used to generate avatar movements.

[0207] The server generates an avatar movement sequence based on this sign language gesture data, which specifically instructs the movement of each joint and body part of the avatar.The server then analyzes the music file and combines a timeline with the rhythm to synchronize the avatar movement sequence with the rhythm and tone of the music.

[0208] Furthermore, the server uses an emotion engine to recognize the user's emotions. The emotion engine uses facial recognition and voice analysis to detect the user's emotional state in real time. The recognized emotion information is reflected in the avatar's movements and sign language dance.

[0209] For example, if the user is expressing a happy emotion, the avatar's motion sequence may include more lively and joyful movements, whereas if the user is sad, the avatar may be controlled to include corresponding emotional movements.

[0210] The server then generates a dance video of the avatar based on the movement sequence. The video is visually easy to understand, with the avatar's movements and music perfectly matched. The video is saved in a common video format, such as MP4.

[0211] The server then distributes the generated video over the network, uploading the generated video file to a social networking platform, such as a video content sharing application.

[0212] Finally, users can watch the video on social media platforms. The video is presented in a short format, making it easy to watch and share. The video also includes relevant metadata (e.g., original lyrics, sign language meanings, etc.), making it easy for viewers to understand the sign language content.

[0213] For example, if a user inputs the lyrics "Spring, Come" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on the data, generates a dance video of the avatar in time with the music rhythm, and finally distributes it to a social media platform. During this process, the emotion engine recognizes the user's emotions in real time and adds corresponding emotional expressions to the avatar's movements.

[0214] The implementation of this system will enable sign language learning and provide a musical experience for the hearing-blind, and by adding actions that match the user's emotions, it is expected to further increase interest in sign language.

[0215] The processing flow will be explained below.

[0216] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[0217] The process flow will be explained in detail below for each step.

[0218] Step 1:

[0219] The user inputs the lyrics of a song into the system. For example, the user can copy the lyrics from a music streaming service app and paste them into the system's dedicated input form. This step begins by clicking the "Convert to Sign Language" button.

[0220] Step 2:

[0221] The device sends lyric text data to the server, converts the input lyric data into a JSON-formatted request, and sends it to the server. This communication is carried out via a secure protocol such as HTTPS.

[0222] Step 3:

[0223] The server receives the lyrics text and inputs it into the AI ​​model. The server analyzes the lyrics text and issues a request to the AI ​​model's API. The AI ​​model analyzes the lyrics and generates corresponding sign language gesture data.

[0224] Step 4:

[0225] The server receives the generated sign language data and stores it in an internal database, which is used later to generate avatar movements.

[0226] Step 5:

[0227] The server generates an avatar movement sequence from the sign language data.The server generates a movement sequence that specifically defines the joint and body movements of the avatar based on the sign language gesture data.

[0228] Step 6:

[0229] The server recognizes the user's emotions using an emotion engine. The emotion engine uses facial recognition and voice analysis of the user to detect the user's emotional state in real time. The emotion engine analyzes data input from the user's camera and microphone to identify whether the user is happy, sad, or other emotions.

[0230] Step 7:

[0231] The server adds emotional expressions to the avatar's motion sequence based on the recognized emotion information. For example, if the user is expressing a sense of enjoyment, the server adds more lively and joyful gestures to the avatar's motion. Conversely, if the user is sad, the server includes corresponding sadness-expressing motions.

[0232] Step 8:

[0233] The server generates a dance video of the avatar based on the sign language sequence, synchronized to the rhythm and tone of the music. The server analyzes the music file and calculates the timing based on the rhythm and tone. The server then matches the sign language sequence to the avatar's movements and generates the animation using 3D modeling software.

[0234] Step 9:

[0235] The server stores the generated dance video data and prepares it for distribution. The server saves the generated dance video in a format such as MP4. The server then converts the video file into a format suitable for distribution on social media and adds metadata (e.g., original lyrics, sign language explanations).

[0236] Step 10:

[0237] The server uploads the created dance video to the SNS platform. The server then sends the saved video file to the SNS platform's API for uploading.

[0238] Step 11:

[0239] Allow users to watch videos on the social media platform, which then makes the uploaded videos public and available for users to watch.

[0240] Step 12:

[0241] To make it easier for users to watch, the videos are distributed in short video format and related metadata (e.g., original lyrics, meaning of sign language) is also provided. The server automatically adds the video file metadata to the video description section of the social media platform.

[0242] For example, if a user inputs the lyrics "Spring, Come" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on the data, generates a dance video of the avatar in time with the music rhythm, and finally distributes it to a social media platform. During this process, the emotion engine recognizes the user's emotions in real time and adds corresponding emotional expressions to the avatar's movements.

[0243] The implementation of this system will enable sign language learning and provide a musical experience for the hearing-blind, and by adding actions that match the user's emotions, it is expected to further increase interest in sign language.

[0244] Example 2

[0245] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0246] Conventional systems that convert music lyrics into sign language simply convert lyrics into sign language and assign actions to an avatar, but do not take into account the user's emotions. As a result, the avatar's movements are monotonous and lack emotional appeal, and they are unable to sufficiently move or empathize with the viewer. There is also a need to provide a richer music experience for people with hearing and visual impairments.

[0247] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0248] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for synchronizing the generated avatar movements with the rhythm of the music to generate video, means for distributing the generated video via a network, and means for recognizing a user's emotions and adding emotional expressions to the avatar movements. This allows the avatar to be given movements and expressions that match the user's emotions, making it possible to provide video content that is more moving and resonates with viewers.

[0249] "Means for converting music lyrics into sign language" refers to a system or program for analyzing music lyrics text and converting it into sign language gesture data.

[0250] The "means for generating avatar movements based on sign language" refers to a script or algorithm for analyzing the generated sign language gesture data and generating a movement sequence for the avatar based thereon.

[0251] "Means for generating a video by matching the generated avatar's movements to the rhythm of music" refers to software or tools for adjusting the avatar's movement sequence to the rhythm or tone of music and generating it as a video.

[0252] "Means for distributing the generated video via a network" refers to a system or platform for distributing the generated video content via the Internet.

[0253] "Means for recognizing user emotions and adding emotional expressions to avatar movements" refers to an engine or algorithm that analyzes user emotions in real time and adds emotional elements to avatar movements based on that analysis.

[0254] A "generative AI model" is a model that uses machine learning and artificial intelligence techniques to convert input data (e.g., lyric text) into another format (e.g., sign language gesture data).

[0255] A "prompt" refers to a sentence or piece of data used as input for an AI model, and often includes specific instructions or questions.

[0256] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[0257] First, the user enters the lyrics of the music into the system. They copy the lyrics from the music distribution service app and paste them into the system's dedicated input form. The process begins by clicking the "Sign Language Conversion" button. The device used by the user is typically a smartphone or PC, and they access the system through a web browser.

[0258] Next, the device sends the lyric text data to the server. The device converts the input lyric data into a JSON format request and sends it to the server. The data is sent to a cloud server (e.g., Azure, AWS) using the HTTP protocol.

[0259] The server analyzes the received lyrics text and inputs it into a generative AI model. The software used includes natural language processing libraries (e.g., spaCy, NLTK) and machine learning frameworks (e.g., TensorFlow, PyTorch). The generative AI model analyzes the lyrics text and generates corresponding sign language gesture data. The generated sign language gesture data contains information that specifically instructs the movement of each joint and body part of the avatar.

[0260] The server generates an avatar's movement sequence based on this sign language gesture data. 3D content creation tools such as Blender and Unity are used to generate the 3D avatar's movements. A music analysis library (e.g., libROSA) is also used to analyze the music file and combine the movement sequence into a timeline in accordance with the rhythm and tone of the music.

[0261] Furthermore, the server uses an emotion engine to recognize the user's emotions. The emotion engine uses facial recognition technology (e.g., OpenCV, dlib) and voice analysis libraries (e.g., Google Cloud Speech-to-Text API) to detect the user's emotional state in real time. The recognized emotional information is reflected in the avatar's movements and sign language dance. For example, if the user is showing signs of enjoyment, a movement expressing joy will be added to the avatar's movement sequence.

[0262] The server generates a dance video of the avatar based on the movement sequence. It uses Blender or Unity to render the avatar's movements, and then uses FFmpeg to convert and save them as standard video files such as MP4.

[0263] The server then distributes the generated video over the network, and uploads the generated video file to social media platforms such as YouTube and Instagram, making it easily accessible to users and viewers.

[0264] Finally, users can watch the videos on social media platforms. The videos are presented in a short format, making them easy to watch and share. The videos also include relevant metadata, such as the original lyrics and sign language meanings, making it easy for viewers to understand the sign language content.

[0265] Here is a specific example: When a user inputs the lyrics "Haru yo, Koi" (Come Spring) and converts it into sign language, the process proceeds as follows:

[0266] Example prompt:

[0267] 1. The user accesses the system using a web browser, pastes the lyrics of "Haru yo, Koi" into the input form, and clicks the "Convert to Sign Language" button.

[0268] 2. The device converts the entered lyrics into a JSON format request and sends it to the server using the HTTP protocol.

[0269] 3. The server analyzes the lyrics text and generates sign language gesture data using a generative AI model.

[0270] 4. The server generates a movement sequence for the avatar based on the sign language gesture data, and the emotion engine recognizes the user's emotions and reflects them in the movement.

[0271] 5. The server generates the final avatar dance video and uploads it to a social media platform.

[0272] 6. The user watches the video on a social media platform and checks the meaning of the sign language.

[0273] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0274] Step 1:

[0275] The user enters the lyrics text.

[0276] Users access the system through a web browser, paste the music lyrics into the input form, and then click the "Convert to Sign Language" button.

[0277] Input: Lyric text (e.g. "Haru yo, Koi" (Come Spring))

[0278] Output: Signals the start of a request to the system

[0279] Step 2:

[0280] The terminal transmits the lyrics text data to the server.

[0281] The device receives the lyrics data entered by the user, converts it into a JSON-formatted request, and then sends the request to the server using the HTTP protocol.

[0282] Input: Lyric text

[0283] Data processing: Convert lyrics data into JSON format

[0284] Output: HTTP request to the server

[0285] Step 3:

[0286] The server parses the lyrics text.

[0287] The server parses the received JSON request and extracts the lyrics text, then uses a natural language processing library (e.g., spaCy) to analyze the lyrics and extract important keywords and phrases.

[0288] Input: Lyrics data in JSON format

[0289] Data processing: Lyric text analysis, keyword extraction

[0290] Output: Analyzed text data (keywords and phrases)

[0291] Step 4:

[0292] The server generates sign language gesture data.

[0293] The server uses a generative AI model (e.g., a TensorFlow model) to convert the parsed text data into sign language gesture data.

[0294] Input: Parsed text data

[0295] Data Computation: Transformation with Generative AI Models

[0296] Output: Sign language gesture data

[0297] Step 5:

[0298] The server generates the avatar movement sequence.

[0299] The server generates avatar movement sequences based on sign language gesture data using 3D content creation tools such as Blender and Unity, and analyzes music files to combine movement sequences into a timeline in sync with the rhythm and tone of the music.

[0300] Input: Sign language gesture data, music files

[0301] Data processing: generating motion sequences and combining timelines

[0302] Output: Avatar movement sequence

[0303] Step 6:

[0304] The server uses an emotion engine to recognize the user's emotion.

[0305] The server uses facial recognition technology (e.g., OpenCV, dlib) and voice analysis libraries (e.g., Google Cloud Speech-to-Text API) to detect the user's emotional state in real time. The recognized emotional information is reflected in the avatar's behavior.

[0306] Input: User's face image, voice data

[0307] Data Computing: Applying Emotion Recognition Algorithms

[0308] Output: User's emotional state information

[0309] Step 7:

[0310] The server generates a dance video of the avatar.

[0311] The server generates a dance video of the avatar based on the final movement sequence. It renders the avatar's movements using Blender or Unity and exports them to a video file (e.g., MP4 format) using FFmpeg.

[0312] Input: Action sequence, emotional state information

[0313] Data processing: video rendering, file export

[0314] Output: Dance video file

[0315] Step 8:

[0316] The server distributes the generated video over the network.

[0317] The server uploads the generated video file to a social media platform (e.g. YouTube, Instagram), allowing users to access the video file.

[0318] Input: Dance video file

[0319] Data transmission: Video upload

[0320] Output: Public videos on social media platforms

[0321] Step 9:

[0322] A user watches a video on a social media platform.

[0323] Users access the social media platform and watch videos of the generated avatar dancing, with associated metadata such as the original lyrics and sign language meanings displayed alongside the video.

[0324] Input: Access to social media platforms

[0325] Output: Watchable dance video and associated metadata

[0326] (Application example 2)

[0327] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0328] Conventional sign language translation systems cannot reflect the user's emotions when converting music lyrics into sign language and expressing them with an avatar, making it difficult to add emotional expression to the sign language movements or avatar dance. Furthermore, even if the generated video content is visually rich, it lacks movements that correspond to the user's emotions, limiting the viewing experience.

[0329] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0330] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for generating a video by synchronizing the generated avatar movements with the rhythm of the music, means for recognizing the user's emotions and adding emotional expressions to the avatar movements, and means for distributing the generated video via a network. This makes it possible to generate a sign language dance video that incorporates expressions according to the user's emotions, providing a richer viewing experience.

[0331] "Music" is an art form that expresses human emotion and experience in the form of sound, and includes rhythm, melody, and harmony.

[0332] "Lyrics" are a part of music that expresses the content and theme of a song in words.

[0333] "Sign language" is a visual, gestural language used for communication between people with hearing impairments, using hand and finger movements and facial expressions.

[0334] An "avatar" refers to a virtual person or character that symbolically represents a user or character in a digital space.

[0335] "Movement" refers to the physical movements and gestures performed by the avatar, including things like sign language and dancing.

[0336] "Video" is a media format that visually recreates movement by displaying a series of still images in succession.

[0337] "Rhythm" refers to the pattern of changes in the length and dynamics of notes in music, and is an element that gives a piece of music a sense of unity and movement.

[0338] "Emotions" refer to the psychological states that humans feel in response to specific situations or events, and include joy, sadness, anger, etc.

[0339] A "network" is a system of computers and communication devices connected together to share information and data.

[0340] "Distribution" refers to the act of providing or transmitting digital content to users over a network.

[0341] A "server" is a computer system that provides data and services over a network.

[0342] A "generative AI model" is a model generated by artificial intelligence and trained to predict or generate appropriate outputs for specific inputs.

[0343] A "prompt sentence" is an input sentence provided to a generative AI model to obtain a desired output.

[0344] This system converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates a function to recognize the user's emotions and add emotional expressions to the avatar's movements.

[0345] System Configuration

[0346] The system consists of the following main components: a user device, a server, a generative AI model, an emotion recognition engine, and a video distribution platform.

[0347] Program processing description

[0348] 1. User Device

[0349] - Lyrics text input: The user device provides an input form for entering music lyrics. The user enters the lyrics text of their favorite song into this form and clicks the "Convert to sign language" button.

[0350] - Emotion recognition: The user's face is captured in real time through the camera on the user's device, and the emotion recognition engine analyzes the user's emotions.

[0351] 2. Data Transmission

[0352] - The device converts the entered lyrics data into JSON format and sends it to the server. At the same time, the user's emotional data is also sent.

[0353] 3. Server Processing

[0354] - The server analyzes the received lyrics text and converts them into sign language gesture data using a generative AI model (e.g., OpenAI's GPT-4). This conversion process uses prompts.

[0355] - Use an emotion recognition engine (e.g., Microsoft Azure Emotion Recognition API) to obtain the user's emotional information and reflect it in the avatar's behavior.

[0356] 4. Avatar Motion Generation

[0357] The server generates a movement sequence for the avatar based on the generated sign language gesture data, which specifically instructs the movements of each joint and body part of the avatar.

[0358] - Perform rhythmic analysis of the music (e.g., librosa library) and synchronize the avatar's movement sequences to the music.

[0359] 5. Video Generation

[0360] - The server generates a dance video by combining the avatar's movement sequence with music, using a video generation tool (e.g., OpenCV and ffmpeg library). The generated video is saved in MP4 format.

[0361] 6. Video streaming

[0362] - The server distributes the generated videos over the network, and uploads them to social media platforms (e.g. YouTube API, Facebook API) so that users can easily watch and share them within the app.

[0363] Examples of concrete examples and prompts

[0364] As a concrete example, when the lyrics "Haru yo, Koi" (Come Spring) are entered by a user, the generative AI model generates sign language gesture data using the following prompt sentence:

[0365] Input lyrics:

[0366] "Come Spring"

[0367] Prompt statement:

[0368] Convert the following lyrics into sign language and generate sign language gesture data for the avatar. Also, please consider that the user's emotion is "joy." Output the generated gesture data in JSON format.

[0369] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0370] Step 1:

[0371] The user inputs music lyric text into the terminal and clicks the "Convert to Sign Language" button.

[0372] Input: Music lyric text

[0373] Output: Lyric text input confirmation

[0374] First, the user enters the lyrics of the music into the application's dedicated input form and clicks the "Convert to Sign Language" button. At this time, the user's camera is activated and captures facial images in real time, which prepares the emotion recognition engine to analyze the user's emotions.

[0375] Step 2:

[0376] The device converts the lyrics text into JSON format and sends it to the server, along with the user's facial image data.

[0377] Input: Lyric text, user's facial image data

[0378] Output: Sending lyrics data in JSON format and facial video data to the server

[0379] The device converts the entered lyrics text into JSON format and sends it to a cloud server. A video of the user's face is also transferred at the same time, allowing subsequent processing to proceed smoothly.

[0380] Step 3:

[0381] The server analyzes the lyric text and generates sign language gesture data using a generative AI model.

[0382] Input: Lyrics data in JSON format

[0383] Output: Sign language gesture data

[0384] The server parses the received JSON-formatted lyrics data and converts the lyrics text into sign language gesture data using a generative AI model (e.g., OpenAI's GPT-4). During this process, an appropriate prompt is passed to the generative AI model.

[0385] Step 4:

[0386] The server uses an emotion recognition engine to analyze the user's emotions and obtain emotion data.

[0387] Input: Facial image data

[0388] Output: Emotion data

[0389] The server uses an emotion recognition engine (e.g., Microsoft Azure Emotion Recognition API) to analyze the facial video data and identify the user's emotions. The emotion data is saved so that it can be reflected in the avatar's behavior.

[0390] Step 5:

[0391] The server generates an avatar's movement sequence based on the sign language gesture data and emotion data.

[0392] Input: Sign language gesture data, emotion data

[0393] Output: Avatar movement sequence

[0394] The server combines the sign language gesture data and emotion data to generate a movement sequence for the avatar, which generates realistic and emotional avatar movements based on the user's emotions. Specific instructions for the movements of each joint and body part are created.

[0395] Step 6:

[0396] The server analyzes the rhythm of the music and synchronizes the avatar's movement sequence with the music.

[0397] Input: Action sequences, music files

[0398] Output: rhythm-synchronized movement sequence

[0399] The server analyzes the music files to identify the rhythm (e.g., using the librosa library), then combines the rhythmic timeline with a movement sequence to create engaging, music-synchronized avatar movements.

[0400] Step 7:

[0401] The server combines the avatar's movement sequence with music to generate a video.

[0402] Input: rhythm-synchronized movement sequences, music files

[0403] Output: Generated video file

[0404] The server finally generates a dance video by combining the avatar's movement sequence with music. This video generation uses libraries such as OpenCV and ffmpeg to generate a video file in MP4 format.

[0405] Step 8:

[0406] The server distributes the generated video to the social media platform via the network.

[0407] Input: Generated video file

[0408] Output: Upload to social media platforms

[0409] The server uploads the generated video to a social networking platform (e.g., YouTube API, Facebook API) for users to easily access, watch, and share, allowing users and viewers to enjoy the emotional avatar movements and music.

[0410] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0411] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0412] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0413] [Second embodiment]

[0414] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0415] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0416] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0417] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0418] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0419] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0420] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0421] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0422] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0423] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0424] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0425] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0426] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing in accompaniment with the sign language. This system performs a series of processes, including inputting lyrics text, converting them into sign language, generating avatar movements, generating videos, and distributing them.

[0427] First, the user inputs the lyrics of the music into the system using their own device. For example, the user can copy the lyrics from a music distribution service app and paste them into the system's dedicated input form. Then, the process begins by clicking the "Sign Language Conversion" button.

[0428] Next, the device sends the lyrics text data to the server, where it converts the input lyrics data into JSON format and sends it to the server using a dedicated communication protocol for fast and reliable data transfer.

[0429] The server analyzes the received lyrics text and inputs it into a pre-trained AI model, which then analyzes the lyrics text and generates corresponding sign language gesture data, which is then used to generate avatar movements.

[0430] The server generates an avatar movement sequence based on this sign language gesture data, which specifically instructs the movement of each joint and body part of the avatar.The server then analyzes the music file to synchronize the avatar movement sequence with the rhythm and tone of the music, and combines a timeline with the rhythm.

[0431] The server then generates a dance video of the avatar based on the movement sequence, matching the avatar's movements with the music to create a visually appealing video. The video is saved in a common video format, such as MP4.

[0432] The server then distributes the generated video over the network, uploading the video file to a dedicated social networking platform, such as a video content sharing application.

[0433] Finally, users can watch the video on social media platforms. The video is presented in a short format, making it easy to watch and share. The video also includes relevant metadata (e.g., original lyrics, sign language meanings, etc.), making it easy for viewers to understand the sign language content.

[0434] As a concrete example, if a user inputs the lyrics "Come Spring" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on that, generates a dance video of the avatar in time with the rhythm of the music, and finally distributes it to a social media platform.

[0435] The implementation of this system will enable people with visual and hearing impairments to learn sign language and to experience music, and is expected to increase interest in sign language.

[0436] The processing flow will be explained below.

[0437] Step 1:

[0438] The user inputs the lyrics of a song into the system. The user copies the lyrics from a music distribution service app and pastes them into the system's dedicated input form. Then, the user clicks the "Convert to Sign Language" button.

[0439] Step 2:

[0440] The device sends lyric text data to the server. The device converts the input lyric data into a JSON format request and sends it to the server.

[0441] Step 3:

[0442] The server receives the lyrics text and inputs it into the AI ​​model. The server analyzes the lyrics text and issues a request to the AI ​​model's API. The AI ​​model analyzes the lyrics and generates corresponding sign language gesture data.

[0443] Step 4:

[0444] The server receives the generated sign language data and stores it in an internal database, which is used to generate movement sequences for the avatar.

[0445] Step 5:

[0446] The server generates an avatar's movement sequence from the sign language data. The server executes a script that generates a sequence corresponding to the avatar's joint and body movements based on the sign language gesture data.

[0447] Step 6:

[0448] The server generates a dance video of the avatar based on the sign language sequence and synchronizes it with the rhythm and tone of the music. The server analyzes the music file and calculates the timing based on the rhythm and tone. The server then reflects the sign language sequence in the avatar's movements according to the timing and generates the animation using 3D modeling software.

[0449] Step 7:

[0450] The server saves the generated dance video data and prepares it for distribution. The server saves the generated dance video in a format such as MP4. The server converts the video file into a format suitable for distribution on social media and adds metadata (e.g., original lyrics, sign language explanations).

[0451] Step 8:

[0452] The server uploads the created dance video to the SNS platform. The server then sends the saved video file to the SNS platform's API for uploading.

[0453] Step 9:

[0454] Allow users to watch videos on the social media platform, which then makes the uploaded videos public and available for users to watch.

[0455] Step 10:

[0456] To make it easier for users to watch, the videos are distributed in short video format and related metadata (e.g., original lyrics, meaning of sign language) is also provided. The server automatically adds the video file metadata to the video description section of the social media platform.

[0457] Example 1

[0458] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0459] With conventional technology, it has been difficult to convert music lyrics into sign language and efficiently generate and distribute videos of avatars dancing using sign language. Furthermore, there have been limited ways for the hearing-blind to enjoy music, and there has also been a lack of methods to promote learning sign language. This has limited musical experiences and learning opportunities for them.

[0460] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0461] In this invention, the server includes means for converting lyric text data into JSON format and sending it to the server, means for analyzing the lyric text data using an AI model and converting it into sign language gesture data, and means for generating an avatar movement sequence based on the sign language gesture data. This allows the generation of an avatar movement sequence in sync with the rhythm of music, providing a means for the hearing-blind to enjoy music and promoting the learning of sign language.

[0462] "Lyric text data" refers to data in which music lyrics are saved as text information.

[0463] "Sign language gesture data" is data that expresses each gesture or movement in sign language.

[0464] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format that consists of data in key-value pairs.

[0465] An "AI model" is an algorithm or system trained using artificial intelligence techniques to perform a specific task or analysis.

[0466] An "avatar" refers to a virtual person or character in a digital environment.

[0467] A "movement sequence" is a series of commands that specify the movement of each joint and body part of the avatar along a time axis.

[0468] "3D modeling software" is software for creating, editing, and displaying three-dimensional models.

[0469] The "MP4 format" is a common digital file format for storing video and audio.

[0470] A "network" is a system of computers and other devices connected to exchange data with each other.

[0471] An "SNS platform" is an online service that allows users to post, share, and view content.

[0472] "Sign language" is a system of visual gestures and movements that enables deaf people to communicate.

[0473] The present invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing in accompaniment with the sign language. An embodiment of this system will be described in detail below.

[0474] First, the user inputs the lyrics of the music using their own device. The user copies the lyrics from a music distribution service application or similar and pastes them into the designated input form. Then, the process begins by clicking the "Sign Language Conversion" button.

[0475] Next, the device converts the entered lyrics text data into JSON format and sends it to the server. The device uses a dedicated communication protocol for this data conversion and transmission.

[0476] The server analyzes the received lyric text data using a pre-trained generative AI model, which converts the lyric text into sign language gesture data, which is then used to generate a movement sequence for the avatar.

[0477] The server then generates a motion sequence for the avatar based on the sign language gesture data, which specifically instructs the movements of each joint and body part of the avatar. It also analyzes the music file and combines a timeline with the rhythm of the music to synchronize the avatar's movements with the rhythm of the music.

[0478] The server then generates a dance video of the avatar based on the movement sequence, using 3D modeling software. The video, which matches the avatar's movements with the music, is visually appealing and saved in a common video format such as MP4.

[0479] Finally, the server distributes the generated video over the network. The video file is uploaded to a dedicated social networking platform, where users can watch the video. The video also includes relevant metadata, such as the original lyrics and sign language meanings, making it easy for viewers to understand the sign language content.

[0480] For example, if a user inputs the lyrics "Spring Comes" into the system and clicks the "Sign Language Conversion" button, the server converts the lyrics into sign language and generates sign language gesture data. Next, based on the generated sign language gesture data, the server creates a movement sequence for the avatar and generates a video of the avatar dancing to the rhythm of the music. Finally, this video is distributed to a social networking platform.

[0481] Example prompt sentence:

[0482] "Enter the following lyrics, 'Spring Comes,' into the system, convert them into sign language, generate a dance video with an avatar, and post it on social media."

[0483] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0484] Step 1:

[0485] A user uses his / her terminal to input the lyrics text of the music.

[0486] Specific operation: The user copies lyrics from a music distribution service application, pastes them into a dedicated input form, and clicks the "Convert to sign language" button.

[0487] Input: Music lyric text

[0488] Output: Trigger the start of processing by clicking the "Sign Language Conversion" button

[0489] Step 2:

[0490] The device converts the entered lyrics text data into JSON format and sends it to the server.

[0491] Specific operation: The terminal converts the text data into JSON format and sends it to the server using a dedicated communication protocol.

[0492] Input: Lyric text entered by the user

[0493] Output: Lyrics text data in JSON format sent to the server

[0494] Step 3:

[0495] The server analyzes the received lyrics text data.

[0496] Specific operation: The server analyzes the received JSON format lyrics text data using the analysis module.

[0497] Input: Lyrics text data in JSON format

[0498] Output: Parsed lyrics text data

[0499] Step 4:

[0500] The server inputs the analyzed lyrics text data into a generative AI model to generate sign language gesture data.

[0501] Specific operation: Text data is input into the generative AI model, and sign language gesture data is generated.

[0502] Input: Parsed lyrics text data

[0503] Output: Sign language gesture data

[0504] Step 5:

[0505] The server generates a movement sequence for the avatar based on the sign language gesture data.

[0506] Specific movements: Based on sign language gesture data, a movement sequence is generated that specifically instructs the movement of each joint and body part of the avatar.

[0507] Input: Sign language gesture data

[0508] Output: Avatar movement sequence

[0509] Step 6:

[0510] The server analyzes the music files and combines the movement sequences with rhythm.

[0511] Specific Actions: Analyze music files and combine action sequences into a timeline based on the rhythm of the music.

[0512] Input: Avatar movement sequence, music file

[0513] Output: Motion sequence synchronized with the music rhythm

[0514] Step 7:

[0515] The server generates a dance video based on the movement sequence.

[0516] Specific Movement: Using 3D modeling software, a dance video is generated that combines an avatar's movement sequence with music.

[0517] Input: A sequence of movements synchronized with the musical rhythm

[0518] Output: Generated dance video (MP4 format, etc.)

[0519] Step 8:

[0520] The server distributes the generated video over the network.

[0521] Specific operation: Upload the generated video file to a social media platform.

[0522] Input: Generated dance video (e.g. MP4 format)

[0523] Output: Upload completion notification to SNS platform

[0524] Step 9:

[0525] Users watch videos on social media platforms.

[0526] Specific Action: A user visits a social media platform to watch and share a video.

[0527] Input: Video link on social media platform

[0528] Output: Videos watched by users, videos shared

[0529] (Application example 1)

[0530] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0531] Conventional food delivery services have the problem that it is difficult for visually or hearing impaired users to understand the details of the dishes. In addition, there are many cases where explanations are not available in sign language, and a sign language interpreter is required. There is a need for a means to enable visually or audibly impaired users to understand detailed descriptions of dishes.

[0532] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0533] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for generating a video in which the generated avatar movements are synchronized with the rhythm of the music, means for distributing the generated video via a network, means for converting text describing dishes into sign language and generating a video in which an avatar uses sign language, and means for distributing the generated sign language video to visually impaired users, thereby enabling visually or hearing impaired users to understand the details of dishes when using a food delivery service.

[0534] "Musical lyrics" are text, usually consisting of words, that is sung to accompany a melody in a musical performance or song.

[0535] "Sign language" is a visual language that primarily communicates through hand and finger movements and facial expressions.

[0536] An "avatar" is a character or graphical representation that represents a user and acts on behalf of the user in a virtual space or on a computer.

[0537] "Generating behavior" refers to creating a series of actions or poses for an avatar or character on a computer.

[0538] "Musical rhythm" refers to the temporal pattern or tempo of music.

[0539] "Generating video" refers to creating a moving image by displaying a series of images or frames in succession.

[0540] "Distributing" refers to transmitting digital content over a network and making the content viewable on the receiving end.

[0541] "Description text of a dish" is a text about the contents, characteristics, ingredients, cooking method, etc. of a dish.

[0542] "Sign language gesture data" is a digital representation of sign language movements, and is typically data that indicates the position, direction, and movement of the hands and fingers.

[0543] "Visually impaired people" refers to people who have impaired vision and who have difficulty or are unable to understand visual information.

[0544] "Hearing impaired people" refers to people who have hearing impairments and who have difficulty or are unable to understand auditory information.

[0545] An "AI model" refers to an algorithm or computational model built using machine learning or deep learning, which performs data analysis and predictions.

[0546] "Via a network" refers to sending and receiving data via a communications infrastructure such as the Internet or a local area network.

[0547] This invention provides a system that enables visually and hearing impaired users to understand the details of dishes in food delivery services. Specifically, we apply technology that converts music lyrics into sign language to build a system that converts text describing dishes into sign language and generates and distributes videos using sign language by an avatar.

[0548] Basic system configuration

[0549] The server has the following functions:

[0550] 1. Text analysis:

[0551] The system analyzes the text description of a dish entered by the user and converts it into sign language gesture data using a generative AI model (using PyTorch and TensorFlow).

[0552] 2. Avatar movement generation:

[0553] Based on the analyzed sign language gesture data, a movement sequence for the avatar is generated, which specifically specifies the movements of each joint and body part required for sign language.

[0554] 3. Speech generation:

[0555] Use Google Cloud Speech-to-Text and Amazon Polly to convert the text of the cooking instructions into an audio file and add audio to the sign language video.

[0556] 4. Video Generation:

[0557] Using a video editing API such as OpenCV, sign language videos are generated so that the avatar's movements match the rhythm of the music.

[0558] 5. Delivery:

[0559] Using AWS S3 and CloudFront, the generated sign language videos are distributed to visually and hearing impaired people within a food delivery service application.

[0560] The user terminal has the following functions:

[0561] 1. Input interface:

[0562] It provides an interface for inputting text describing a dish. The user inputs the text and clicks the "Convert to sign language" button to start the process.

[0563] 2. Watching videos:

[0564] The distributed sign language video is played within the application, allowing the content to be conveyed to visually and hearing impaired people.

[0565] Processing flow

[0566] The user enters a text description of the dish: "This curry is made with special spices and is mild." The user's device sends the entered text data to the server, which converts the text into sign language gesture data. Next, the avatar movement generation function generates an avatar movement sequence based on the sign language gesture data. Furthermore, the audio file generated using Google Cloud Speech-to-Text and Amazon Polly is used to generate a sign language video of the avatar using a video editing API such as OpenCV. Finally, this video is distributed via AWS S3 and CloudFront and played on the user's device.

[0567] Prompt Sentence Examples

[0568] Below are some examples of dish descriptions:

[0569] "This curry is made with special spices and is not too spicy."

[0570] This pasta is made with plenty of seasonal vegetables.

[0571] This will enable users with visual or hearing impairments to understand details about dishes through food delivery services.

[0572] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0573] Step 1: User Input

[0574] The user inputs a text description of a dish through the food delivery service's application. For example, they input "This curry is made with special spices and is mild," and then click the "Sign Language Conversion" button. The input is sent to the device as text data.

[0575] Step 2: Send text data

[0576] The terminal converts the input text data into JSON format and sends it to the server. An appropriate communication protocol is used to ensure accurate data transfer. The input is text data and the output is JSON format data.

[0577] Step 3: Generate sign language gesture data

[0578] The server analyzes the received text data and inputs it into a generative AI model (using PyTorch or TensorFlow). The model converts the text into sign language gesture data. For example, it generates sign language gesture data for "curry." The input is text data, and the output is sign language gesture data.

[0579] Step 4: Generate avatar movement sequences

[0580] The server analyzes the sign language gesture data and generates a movement sequence for the avatar. A script is generated that specifically instructs the movement of each joint and body part, and the avatar's movement is determined based on this. The input is the sign language gesture data, and the output is the avatar's movement sequence.

[0581] Step 5: Speech generation

[0582] The server converts the text of the cooking instructions into an audio file using Google Cloud Speech-to-Text and Amazon Polly. The generated audio file is used to accompany the sign language video. The input is text data, and the output is an audio file.

[0583] Step 6: Video Generation

[0584] The server uses a video editing API such as OpenCV to combine the avatar's movement sequence based on the sign language gesture data with the generated audio file to generate a sign language video. This video is visually easy to understand, with the sign language and audio synchronized. The input is the avatar's movement sequence and audio file, and the output is the final sign language video.

[0585] Step 7: Publish your video

[0586] The server distributes the generated sign language video using AWS S3 and CloudFront, allowing users to watch the video on their devices. The input is the sign language video file, and the output is the video URL distributed over the network.

[0587] Step 8: Watch the video

[0588] The user device plays the sign language video delivered from the server and communicates the details of the dish to the visually and hearing impaired. The user can watch the video within the app and understand the explanation. The input is the video URL, and the output is the sign language video that is played.

[0589] This allows users with visual or hearing impairments to understand the details of the dishes and make appropriate choices.

[0590] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0591] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[0592] A specific embodiment of this system is described below.

[0593] First, the user inputs the lyrics of a song into the system. The user copies the lyrics from the music distribution service app and pastes them into the system's dedicated input form. The process begins by clicking the "Convert to Sign Language" button.

[0594] Next, the device sends the lyrics text data to the server, converts the input lyrics data into a JSON format request, and sends it to the server.

[0595] The server analyzes the received lyrics text and inputs it into the AI ​​model, which then analyzes the lyrics text and generates corresponding sign language gesture data, which is then used to generate avatar movements.

[0596] The server generates an avatar movement sequence based on this sign language gesture data, which specifically instructs the movement of each joint and body part of the avatar.The server then analyzes the music file and combines a timeline with the rhythm to synchronize the avatar movement sequence with the rhythm and tone of the music.

[0597] Furthermore, the server uses an emotion engine to recognize the user's emotions. The emotion engine uses facial recognition and voice analysis to detect the user's emotional state in real time. The recognized emotion information is reflected in the avatar's movements and sign language dance.

[0598] For example, if the user is expressing a happy emotion, the avatar's motion sequence may include more lively and joyful movements, whereas if the user is sad, the avatar may be controlled to include corresponding emotional movements.

[0599] The server then generates a dance video of the avatar based on the movement sequence. The video is visually easy to understand, with the avatar's movements and music perfectly matched. The video is saved in a common video format, such as MP4.

[0600] The server then distributes the generated video over the network, uploading the generated video file to a social networking platform, such as a video content sharing application.

[0601] Finally, users can watch the video on social media platforms. The video is presented in a short format, making it easy to watch and share. The video also includes relevant metadata (e.g., original lyrics, sign language meanings, etc.), making it easy for viewers to understand the sign language content.

[0602] For example, if a user inputs the lyrics "Spring, Come" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on the data, generates a dance video of the avatar in time with the music rhythm, and finally distributes it to a social media platform. During this process, the emotion engine recognizes the user's emotions in real time and adds corresponding emotional expressions to the avatar's movements.

[0603] The implementation of this system will enable sign language learning and provide a musical experience for the hearing-blind, and by adding actions that match the user's emotions, it is expected to further increase interest in sign language.

[0604] The processing flow will be explained below.

[0605] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[0606] The process flow will be explained in detail below for each step.

[0607] Step 1:

[0608] The user inputs the lyrics of a song into the system. For example, the user can copy the lyrics from a music streaming service app and paste them into the system's dedicated input form. This step begins by clicking the "Convert to Sign Language" button.

[0609] Step 2:

[0610] The device sends lyric text data to the server, converts the input lyric data into a JSON-formatted request, and sends it to the server. This communication is carried out via a secure protocol such as HTTPS.

[0611] Step 3:

[0612] The server receives the lyrics text and inputs it into the AI ​​model. The server analyzes the lyrics text and issues a request to the AI ​​model's API. The AI ​​model analyzes the lyrics and generates corresponding sign language gesture data.

[0613] Step 4:

[0614] The server receives the generated sign language data and stores it in an internal database, which is used later to generate avatar movements.

[0615] Step 5:

[0616] The server generates an avatar movement sequence from the sign language data.The server generates a movement sequence that specifically defines the joint and body movements of the avatar based on the sign language gesture data.

[0617] Step 6:

[0618] The server recognizes the user's emotions using an emotion engine. The emotion engine uses facial recognition and voice analysis of the user to detect the user's emotional state in real time. The emotion engine analyzes data input from the user's camera and microphone to identify whether the user is happy, sad, or other emotions.

[0619] Step 7:

[0620] The server adds emotional expressions to the avatar's motion sequence based on the recognized emotion information. For example, if the user is expressing a sense of enjoyment, the server adds more lively and joyful gestures to the avatar's motion. Conversely, if the user is sad, the server includes corresponding sadness-expressing motions.

[0621] Step 8:

[0622] The server generates a dance video of the avatar based on the sign language sequence, synchronized to the rhythm and tone of the music. The server analyzes the music file and calculates the timing based on the rhythm and tone. The server then matches the sign language sequence to the avatar's movements and generates the animation using 3D modeling software.

[0623] Step 9:

[0624] The server stores the generated dance video data and prepares it for distribution. The server saves the generated dance video in a format such as MP4. The server then converts the video file into a format suitable for distribution on social media and adds metadata (e.g., original lyrics, sign language explanations).

[0625] Step 10:

[0626] The server uploads the created dance video to the SNS platform. The server then sends the saved video file to the SNS platform's API for uploading.

[0627] Step 11:

[0628] Allow users to watch videos on the social media platform, which then makes the uploaded videos public and available for users to watch.

[0629] Step 12:

[0630] To make it easier for users to watch, the videos are distributed in short video format and related metadata (e.g., original lyrics, meaning of sign language) is also provided. The server automatically adds the video file metadata to the video description section of the social media platform.

[0631] For example, if a user inputs the lyrics "Spring, Come" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on the data, generates a dance video of the avatar in time with the music rhythm, and finally distributes it to a social media platform. During this process, the emotion engine recognizes the user's emotions in real time and adds corresponding emotional expressions to the avatar's movements.

[0632] The implementation of this system will enable sign language learning and provide a musical experience for the hearing-blind, and by adding actions that match the user's emotions, it is expected to further increase interest in sign language.

[0633] Example 2

[0634] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0635] Conventional systems that convert music lyrics into sign language simply convert lyrics into sign language and assign actions to an avatar, but do not take into account the user's emotions. As a result, the avatar's movements are monotonous and lack emotional appeal, and they are unable to sufficiently move or empathize with the viewer. There is also a need to provide a richer music experience for people with hearing and visual impairments.

[0636] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0637] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for synchronizing the generated avatar movements with the rhythm of the music to generate video, means for distributing the generated video via a network, and means for recognizing a user's emotions and adding emotional expressions to the avatar movements. This allows the avatar to be given movements and expressions that match the user's emotions, making it possible to provide video content that is more moving and resonates with viewers.

[0638] "Means for converting music lyrics into sign language" refers to a system or program for analyzing music lyrics text and converting it into sign language gesture data.

[0639] The "means for generating avatar movements based on sign language" refers to a script or algorithm for analyzing the generated sign language gesture data and generating a movement sequence for the avatar based thereon.

[0640] "Means for generating a video by matching the generated avatar's movements to the rhythm of music" refers to software or tools for adjusting the avatar's movement sequence to the rhythm or tone of music and generating it as a video.

[0641] "Means for distributing the generated video via a network" refers to a system or platform for distributing the generated video content via the Internet.

[0642] "Means for recognizing user emotions and adding emotional expressions to avatar movements" refers to an engine or algorithm that analyzes user emotions in real time and adds emotional elements to avatar movements based on that analysis.

[0643] A "generative AI model" is a model that uses machine learning and artificial intelligence techniques to convert input data (e.g., lyric text) into another format (e.g., sign language gesture data).

[0644] A "prompt" refers to a sentence or piece of data used as input for an AI model, and often includes specific instructions or questions.

[0645] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[0646] First, the user enters the lyrics of the music into the system. They copy the lyrics from the music distribution service app and paste them into the system's dedicated input form. The process begins by clicking the "Sign Language Conversion" button. The device used by the user is typically a smartphone or PC, and they access the system through a web browser.

[0647] Next, the device sends the lyric text data to the server. The device converts the input lyric data into a JSON format request and sends it to the server. The data is sent to a cloud server (e.g., Azure, AWS) using the HTTP protocol.

[0648] The server analyzes the received lyrics text and inputs it into a generative AI model. The software used includes natural language processing libraries (e.g., spaCy, NLTK) and machine learning frameworks (e.g., TensorFlow, PyTorch). The generative AI model analyzes the lyrics text and generates corresponding sign language gesture data. The generated sign language gesture data contains information that specifically instructs the movement of each joint and body part of the avatar.

[0649] The server generates an avatar's movement sequence based on this sign language gesture data. 3D content creation tools such as Blender and Unity are used to generate the 3D avatar's movements. A music analysis library (e.g., libROSA) is also used to analyze the music file and combine the movement sequence into a timeline in accordance with the rhythm and tone of the music.

[0650] Furthermore, the server uses an emotion engine to recognize the user's emotions. The emotion engine uses facial recognition technology (e.g., OpenCV, dlib) and voice analysis libraries (e.g., Google Cloud Speech-to-Text API) to detect the user's emotional state in real time. The recognized emotional information is reflected in the avatar's movements and sign language dance. For example, if the user is showing signs of enjoyment, a movement expressing joy will be added to the avatar's movement sequence.

[0651] The server generates a dance video of the avatar based on the movement sequence. It uses Blender or Unity to render the avatar's movements, and then uses FFmpeg to convert and save them as standard video files such as MP4.

[0652] The server then distributes the generated video over the network, and uploads the generated video file to social media platforms such as YouTube and Instagram, making it easily accessible to users and viewers.

[0653] Finally, users can watch the videos on social media platforms. The videos are presented in a short format, making them easy to watch and share. The videos also include relevant metadata, such as the original lyrics and sign language meanings, making it easy for viewers to understand the sign language content.

[0654] Here is a specific example: When a user inputs the lyrics "Haru yo, Koi" (Come Spring) and converts it into sign language, the process proceeds as follows:

[0655] Example prompt:

[0656] 1. The user accesses the system using a web browser, pastes the lyrics of "Haru yo, Koi" into the input form, and clicks the "Convert to Sign Language" button.

[0657] 2. The device converts the entered lyrics into a JSON format request and sends it to the server using the HTTP protocol.

[0658] 3. The server analyzes the lyrics text and generates sign language gesture data using a generative AI model.

[0659] 4. The server generates a movement sequence for the avatar based on the sign language gesture data, and the emotion engine recognizes the user's emotions and reflects them in the movement.

[0660] 5. The server generates the final avatar dance video and uploads it to a social media platform.

[0661] 6. The user watches the video on a social media platform and checks the meaning of the sign language.

[0662] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0663] Step 1:

[0664] The user enters the lyrics text.

[0665] Users access the system through a web browser, paste the music lyrics into the input form, and then click the "Convert to Sign Language" button.

[0666] Input: Lyric text (e.g. "Haru yo, Koi" (Come Spring))

[0667] Output: Signals the start of a request to the system

[0668] Step 2:

[0669] The terminal transmits the lyrics text data to the server.

[0670] The device receives the lyrics data entered by the user, converts it into a JSON-formatted request, and then sends the request to the server using the HTTP protocol.

[0671] Input: Lyric text

[0672] Data processing: Convert lyrics data into JSON format

[0673] Output: HTTP request to the server

[0674] Step 3:

[0675] The server parses the lyrics text.

[0676] The server parses the received JSON request and extracts the lyrics text, then uses a natural language processing library (e.g., spaCy) to analyze the lyrics and extract important keywords and phrases.

[0677] Input: Lyrics data in JSON format

[0678] Data processing: Lyric text analysis, keyword extraction

[0679] Output: Analyzed text data (keywords and phrases)

[0680] Step 4:

[0681] The server generates sign language gesture data.

[0682] The server uses a generative AI model (e.g., a TensorFlow model) to convert the parsed text data into sign language gesture data.

[0683] Input: Parsed text data

[0684] Data Computation: Transformation with Generative AI Models

[0685] Output: Sign language gesture data

[0686] Step 5:

[0687] The server generates the avatar movement sequence.

[0688] The server generates avatar movement sequences based on sign language gesture data using 3D content creation tools such as Blender and Unity, and analyzes music files to combine movement sequences into a timeline in sync with the rhythm and tone of the music.

[0689] Input: Sign language gesture data, music files

[0690] Data processing: generating motion sequences and combining timelines

[0691] Output: Avatar movement sequence

[0692] Step 6:

[0693] The server uses an emotion engine to recognize the user's emotion.

[0694] The server uses facial recognition technology (e.g., OpenCV, dlib) and voice analysis libraries (e.g., Google Cloud Speech-to-Text API) to detect the user's emotional state in real time. The recognized emotional information is reflected in the avatar's behavior.

[0695] Input: User's face image, voice data

[0696] Data Computing: Applying Emotion Recognition Algorithms

[0697] Output: User's emotional state information

[0698] Step 7:

[0699] The server generates a dance video of the avatar.

[0700] The server generates a dance video of the avatar based on the final movement sequence. It renders the avatar's movements using Blender or Unity and exports them to a video file (e.g., MP4 format) using FFmpeg.

[0701] Input: Action sequence, emotional state information

[0702] Data processing: video rendering, file export

[0703] Output: Dance video file

[0704] Step 8:

[0705] The server distributes the generated video over the network.

[0706] The server uploads the generated video file to a social media platform (e.g. YouTube, Instagram), allowing users to access the video file.

[0707] Input: Dance video file

[0708] Data transmission: Video upload

[0709] Output: Public videos on social media platforms

[0710] Step 9:

[0711] A user watches a video on a social media platform.

[0712] Users access the social media platform and watch videos of the generated avatar dancing, with associated metadata such as the original lyrics and sign language meanings displayed alongside the video.

[0713] Input: Access to social media platforms

[0714] Output: Watchable dance video and associated metadata

[0715] (Application example 2)

[0716] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0717] Conventional sign language translation systems cannot reflect the user's emotions when converting music lyrics into sign language and expressing them with an avatar, making it difficult to add emotional expression to the sign language movements or avatar dance. Furthermore, even if the generated video content is visually rich, it lacks movements that correspond to the user's emotions, limiting the viewing experience.

[0718] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0719] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for generating a video by synchronizing the generated avatar movements with the rhythm of the music, means for recognizing the user's emotions and adding emotional expressions to the avatar movements, and means for distributing the generated video via a network. This makes it possible to generate a sign language dance video that incorporates expressions according to the user's emotions, providing a richer viewing experience.

[0720] "Music" is an art form that expresses human emotion and experience in the form of sound, and includes rhythm, melody, and harmony.

[0721] "Lyrics" are a part of music that expresses the content and theme of a song in words.

[0722] "Sign language" is a visual, gestural language used for communication between people with hearing impairments, using hand and finger movements and facial expressions.

[0723] An "avatar" refers to a virtual person or character that symbolically represents a user or character in a digital space.

[0724] "Movement" refers to the physical movements and gestures performed by the avatar, including things like sign language and dancing.

[0725] "Video" is a media format that visually recreates movement by displaying a series of still images in succession.

[0726] "Rhythm" refers to the pattern of changes in the length and dynamics of notes in music, and is an element that gives a piece of music a sense of unity and movement.

[0727] "Emotions" refer to the psychological states that humans feel in response to specific situations or events, and include joy, sadness, anger, etc.

[0728] A "network" is a system of computers and communication devices connected together to share information and data.

[0729] "Distribution" refers to the act of providing or transmitting digital content to users over a network.

[0730] A "server" is a computer system that provides data and services over a network.

[0731] A "generative AI model" is a model generated by artificial intelligence and trained to predict or generate appropriate outputs for specific inputs.

[0732] A "prompt sentence" is an input sentence provided to a generative AI model to obtain a desired output.

[0733] This system converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates a function to recognize the user's emotions and add emotional expressions to the avatar's movements.

[0734] System Configuration

[0735] The system consists of the following main components: a user device, a server, a generative AI model, an emotion recognition engine, and a video distribution platform.

[0736] Program processing description

[0737] 1. User Device

[0738] - Lyrics text input: The user device provides an input form for entering music lyrics. The user enters the lyrics text of their favorite song into this form and clicks the "Convert to sign language" button.

[0739] - Emotion recognition: The user's face is captured in real time through the camera on the user's device, and the emotion recognition engine analyzes the user's emotions.

[0740] 2. Data Transmission

[0741] - The device converts the entered lyrics data into JSON format and sends it to the server. At the same time, the user's emotional data is also sent.

[0742] 3. Server Processing

[0743] - The server analyzes the received lyrics text and converts them into sign language gesture data using a generative AI model (e.g., OpenAI's GPT-4). This conversion process uses prompts.

[0744] - Use an emotion recognition engine (e.g., Microsoft Azure Emotion Recognition API) to obtain the user's emotional information and reflect it in the avatar's behavior.

[0745] 4. Avatar Motion Generation

[0746] The server generates a movement sequence for the avatar based on the generated sign language gesture data, which specifically instructs the movements of each joint and body part of the avatar.

[0747] - Perform rhythmic analysis of the music (e.g., librosa library) and synchronize the avatar's movement sequences to the music.

[0748] 5. Video Generation

[0749] - The server generates a dance video by combining the avatar's movement sequence with music, using a video generation tool (e.g., OpenCV and ffmpeg library). The generated video is saved in MP4 format.

[0750] 6. Video streaming

[0751] - The server distributes the generated videos over the network, and uploads them to social media platforms (e.g. YouTube API, Facebook API) so that users can easily watch and share them within the app.

[0752] Examples of concrete examples and prompts

[0753] As a concrete example, when the lyrics "Haru yo, Koi" (Come Spring) are entered by a user, the generative AI model generates sign language gesture data using the following prompt sentence:

[0754] Input lyrics:

[0755] "Come Spring"

[0756] Prompt statement:

[0757] Convert the following lyrics into sign language and generate sign language gesture data for the avatar. Also, please consider that the user's emotion is "joy." Output the generated gesture data in JSON format.

[0758] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0759] Step 1:

[0760] The user inputs music lyric text into the terminal and clicks the "Convert to Sign Language" button.

[0761] Input: Music lyric text

[0762] Output: Lyric text input confirmation

[0763] First, the user enters the lyrics of the music into the application's dedicated input form and clicks the "Convert to Sign Language" button. At this time, the user's camera is activated and captures facial images in real time, which prepares the emotion recognition engine to analyze the user's emotions.

[0764] Step 2:

[0765] The device converts the lyrics text into JSON format and sends it to the server, along with the user's facial image data.

[0766] Input: Lyric text, user's facial image data

[0767] Output: Sending lyrics data in JSON format and facial video data to the server

[0768] The device converts the entered lyrics text into JSON format and sends it to a cloud server. A video of the user's face is also transferred at the same time, allowing subsequent processing to proceed smoothly.

[0769] Step 3:

[0770] The server analyzes the lyric text and generates sign language gesture data using a generative AI model.

[0771] Input: Lyrics data in JSON format

[0772] Output: Sign language gesture data

[0773] The server parses the received JSON-formatted lyrics data and converts the lyrics text into sign language gesture data using a generative AI model (e.g., OpenAI's GPT-4). During this process, an appropriate prompt is passed to the generative AI model.

[0774] Step 4:

[0775] The server uses an emotion recognition engine to analyze the user's emotions and obtain emotion data.

[0776] Input: Facial image data

[0777] Output: Emotion data

[0778] The server uses an emotion recognition engine (e.g., Microsoft Azure Emotion Recognition API) to analyze the facial video data and identify the user's emotions. The emotion data is saved so that it can be reflected in the avatar's behavior.

[0779] Step 5:

[0780] The server generates an avatar's movement sequence based on the sign language gesture data and emotion data.

[0781] Input: Sign language gesture data, emotion data

[0782] Output: Avatar movement sequence

[0783] The server combines the sign language gesture data and emotion data to generate a movement sequence for the avatar, which generates realistic and emotional avatar movements based on the user's emotions. Specific instructions for the movements of each joint and body part are created.

[0784] Step 6:

[0785] The server analyzes the rhythm of the music and synchronizes the avatar's movement sequence with the music.

[0786] Input: Action sequences, music files

[0787] Output: rhythm-synchronized movement sequence

[0788] The server analyzes the music files to identify the rhythm (e.g., using the librosa library), then combines the rhythmic timeline with a movement sequence to create engaging, music-synchronized avatar movements.

[0789] Step 7:

[0790] The server combines the avatar's movement sequence with music to generate a video.

[0791] Input: rhythm-synchronized movement sequences, music files

[0792] Output: Generated video file

[0793] The server finally generates a dance video by combining the avatar's movement sequence with music. This video generation uses libraries such as OpenCV and ffmpeg to generate a video file in MP4 format.

[0794] Step 8:

[0795] The server distributes the generated video to the social media platform via the network.

[0796] Input: Generated video file

[0797] Output: Upload to social media platforms

[0798] The server uploads the generated video to a social networking platform (e.g., YouTube API, Facebook API) for users to easily access, watch, and share, allowing users and viewers to enjoy the emotional avatar movements and music.

[0799] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0800] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0801] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0802] [Third embodiment]

[0803] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0804] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0805] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0806] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0807] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0808] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0809] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0810] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0811] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0812] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0813] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0814] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0815] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing in accompaniment with the sign language. This system performs a series of processes, including inputting lyrics text, converting them into sign language, generating avatar movements, generating videos, and distributing them.

[0816] First, the user inputs the lyrics of the music into the system using their own device. For example, the user can copy the lyrics from a music distribution service app and paste them into the system's dedicated input form. Then, the process begins by clicking the "Sign Language Conversion" button.

[0817] Next, the device sends the lyrics text data to the server, where it converts the input lyrics data into JSON format and sends it to the server using a dedicated communication protocol for fast and reliable data transfer.

[0818] The server analyzes the received lyrics text and inputs it into a pre-trained AI model, which then analyzes the lyrics text and generates corresponding sign language gesture data, which is then used to generate avatar movements.

[0819] The server generates an avatar movement sequence based on this sign language gesture data, which specifically instructs the movement of each joint and body part of the avatar.The server then analyzes the music file to synchronize the avatar movement sequence with the rhythm and tone of the music, and combines a timeline with the rhythm.

[0820] The server then generates a dance video of the avatar based on the movement sequence, matching the avatar's movements with the music to create a visually appealing video. The video is saved in a common video format, such as MP4.

[0821] The server then distributes the generated video over the network, uploading the video file to a dedicated social networking platform, such as a video content sharing application.

[0822] Finally, users can watch the video on social media platforms. The video is presented in a short format, making it easy to watch and share. The video also includes relevant metadata (e.g., original lyrics, sign language meanings, etc.), making it easy for viewers to understand the sign language content.

[0823] As a concrete example, if a user inputs the lyrics "Come Spring" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on that, generates a dance video of the avatar in time with the rhythm of the music, and finally distributes it to a social media platform.

[0824] The implementation of this system will enable people with visual and hearing impairments to learn sign language and to experience music, and is expected to increase interest in sign language.

[0825] The processing flow will be explained below.

[0826] Step 1:

[0827] The user inputs the lyrics of a song into the system. The user copies the lyrics from a music distribution service app and pastes them into the system's dedicated input form. Then, the user clicks the "Convert to Sign Language" button.

[0828] Step 2:

[0829] The device sends lyric text data to the server. The device converts the input lyric data into a JSON format request and sends it to the server.

[0830] Step 3:

[0831] The server receives the lyrics text and inputs it into the AI ​​model. The server analyzes the lyrics text and issues a request to the AI ​​model's API. The AI ​​model analyzes the lyrics and generates corresponding sign language gesture data.

[0832] Step 4:

[0833] The server receives the generated sign language data and stores it in an internal database, which is used to generate movement sequences for the avatar.

[0834] Step 5:

[0835] The server generates an avatar's movement sequence from the sign language data. The server executes a script that generates a sequence corresponding to the avatar's joint and body movements based on the sign language gesture data.

[0836] Step 6:

[0837] The server generates a dance video of the avatar based on the sign language sequence and synchronizes it with the rhythm and tone of the music. The server analyzes the music file and calculates the timing based on the rhythm and tone. The server then reflects the sign language sequence in the avatar's movements according to the timing and generates the animation using 3D modeling software.

[0838] Step 7:

[0839] The server saves the generated dance video data and prepares it for distribution. The server saves the generated dance video in a format such as MP4. The server converts the video file into a format suitable for distribution on social media and adds metadata (e.g., original lyrics, sign language explanations).

[0840] Step 8:

[0841] The server uploads the created dance video to the SNS platform. The server then sends the saved video file to the SNS platform's API for uploading.

[0842] Step 9:

[0843] Allow users to watch videos on the social media platform, which then makes the uploaded videos public and available for users to watch.

[0844] Step 10:

[0845] To make it easier for users to watch, the videos are distributed in short video format and related metadata (e.g., original lyrics, meaning of sign language) is also provided. The server automatically adds the video file metadata to the video description section of the social media platform.

[0846] Example 1

[0847] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0848] With conventional technology, it has been difficult to convert music lyrics into sign language and efficiently generate and distribute videos of avatars dancing using sign language. Furthermore, there have been limited ways for the hearing-blind to enjoy music, and there has also been a lack of methods to promote learning sign language. This has limited musical experiences and learning opportunities for them.

[0849] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0850] In this invention, the server includes means for converting lyric text data into JSON format and sending it to the server, means for analyzing the lyric text data using an AI model and converting it into sign language gesture data, and means for generating an avatar movement sequence based on the sign language gesture data. This allows the generation of an avatar movement sequence in sync with the rhythm of music, providing a means for the hearing-blind to enjoy music and promoting the learning of sign language.

[0851] "Lyric text data" refers to data in which music lyrics are saved as text information.

[0852] "Sign language gesture data" is data that expresses each gesture or movement in sign language.

[0853] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format that consists of data in key-value pairs.

[0854] An "AI model" is an algorithm or system trained using artificial intelligence techniques to perform a specific task or analysis.

[0855] An "avatar" refers to a virtual person or character in a digital environment.

[0856] A "movement sequence" is a series of commands that specify the movement of each joint and body part of the avatar along a time axis.

[0857] "3D modeling software" is software for creating, editing, and displaying three-dimensional models.

[0858] The "MP4 format" is a common digital file format for storing video and audio.

[0859] A "network" is a system of computers and other devices connected to exchange data with each other.

[0860] An "SNS platform" is an online service that allows users to post, share, and view content.

[0861] "Sign language" is a system of visual gestures and movements that enables deaf people to communicate.

[0862] The present invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing in accompaniment with the sign language. An embodiment of this system will be described in detail below.

[0863] First, the user inputs the lyrics of the music using their own device. The user copies the lyrics from a music distribution service application or similar and pastes them into the designated input form. Then, the process begins by clicking the "Sign Language Conversion" button.

[0864] Next, the device converts the entered lyrics text data into JSON format and sends it to the server. The device uses a dedicated communication protocol for this data conversion and transmission.

[0865] The server analyzes the received lyric text data using a pre-trained generative AI model, which converts the lyric text into sign language gesture data, which is then used to generate a movement sequence for the avatar.

[0866] The server then generates a motion sequence for the avatar based on the sign language gesture data, which specifically instructs the movements of each joint and body part of the avatar. It also analyzes the music file and combines a timeline with the rhythm of the music to synchronize the avatar's movements with the rhythm of the music.

[0867] The server then generates a dance video of the avatar based on the movement sequence, using 3D modeling software. The video, which matches the avatar's movements with the music, is visually appealing and saved in a common video format such as MP4.

[0868] Finally, the server distributes the generated video over the network. The video file is uploaded to a dedicated social networking platform, where users can watch the video. The video also includes relevant metadata, such as the original lyrics and sign language meanings, making it easy for viewers to understand the sign language content.

[0869] For example, if a user inputs the lyrics "Spring Comes" into the system and clicks the "Sign Language Conversion" button, the server converts the lyrics into sign language and generates sign language gesture data. Next, based on the generated sign language gesture data, the server creates a movement sequence for the avatar and generates a video of the avatar dancing to the rhythm of the music. Finally, this video is distributed to a social networking platform.

[0870] Example prompt sentence:

[0871] "Enter the following lyrics, 'Spring Comes,' into the system, convert them into sign language, generate a dance video with an avatar, and post it on social media."

[0872] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0873] Step 1:

[0874] A user uses his / her terminal to input the lyrics text of the music.

[0875] Specific operation: The user copies lyrics from a music distribution service application, pastes them into a dedicated input form, and clicks the "Convert to sign language" button.

[0876] Input: Music lyric text

[0877] Output: Trigger the start of processing by clicking the "Sign Language Conversion" button

[0878] Step 2:

[0879] The device converts the entered lyrics text data into JSON format and sends it to the server.

[0880] Specific operation: The terminal converts the text data into JSON format and sends it to the server using a dedicated communication protocol.

[0881] Input: Lyric text entered by the user

[0882] Output: Lyrics text data in JSON format sent to the server

[0883] Step 3:

[0884] The server analyzes the received lyrics text data.

[0885] Specific operation: The server analyzes the received JSON format lyrics text data using the analysis module.

[0886] Input: Lyrics text data in JSON format

[0887] Output: Parsed lyrics text data

[0888] Step 4:

[0889] The server inputs the analyzed lyrics text data into a generative AI model to generate sign language gesture data.

[0890] Specific operation: Text data is input into the generative AI model, and sign language gesture data is generated.

[0891] Input: Parsed lyrics text data

[0892] Output: Sign language gesture data

[0893] Step 5:

[0894] The server generates a movement sequence for the avatar based on the sign language gesture data.

[0895] Specific movements: Based on sign language gesture data, a movement sequence is generated that specifically instructs the movement of each joint and body part of the avatar.

[0896] Input: Sign language gesture data

[0897] Output: Avatar movement sequence

[0898] Step 6:

[0899] The server analyzes the music files and combines the movement sequences with rhythm.

[0900] Specific Actions: Analyze music files and combine action sequences into a timeline based on the rhythm of the music.

[0901] Input: Avatar movement sequence, music file

[0902] Output: Motion sequence synchronized with the music rhythm

[0903] Step 7:

[0904] The server generates a dance video based on the movement sequence.

[0905] Specific Movement: Using 3D modeling software, a dance video is generated that combines an avatar's movement sequence with music.

[0906] Input: A sequence of movements synchronized with the musical rhythm

[0907] Output: Generated dance video (MP4 format, etc.)

[0908] Step 8:

[0909] The server distributes the generated video over the network.

[0910] Specific operation: Upload the generated video file to a social media platform.

[0911] Input: Generated dance video (e.g. MP4 format)

[0912] Output: Upload completion notification to SNS platform

[0913] Step 9:

[0914] Users watch videos on social media platforms.

[0915] Specific Action: A user visits a social media platform to watch and share a video.

[0916] Input: Video link on social media platform

[0917] Output: Videos watched by users, videos shared

[0918] (Application example 1)

[0919] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0920] Conventional food delivery services have the problem that it is difficult for visually or hearing impaired users to understand the details of the dishes. In addition, there are many cases where explanations are not available in sign language, and a sign language interpreter is required. There is a need for a means to enable visually or audibly impaired users to understand detailed descriptions of dishes.

[0921] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0922] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for generating a video in which the generated avatar movements are synchronized with the rhythm of the music, means for distributing the generated video via a network, means for converting text describing dishes into sign language and generating a video in which an avatar uses sign language, and means for distributing the generated sign language video to visually impaired users, thereby enabling visually or hearing impaired users to understand the details of dishes when using a food delivery service.

[0923] "Musical lyrics" are text, usually consisting of words, that is sung to accompany a melody in a musical performance or song.

[0924] "Sign language" is a visual language that primarily communicates through hand and finger movements and facial expressions.

[0925] An "avatar" is a character or graphical representation that represents a user and acts on behalf of the user in a virtual space or on a computer.

[0926] "Generating behavior" refers to creating a series of actions or poses for an avatar or character on a computer.

[0927] "Musical rhythm" refers to the temporal pattern or tempo of music.

[0928] "Generating video" refers to creating a moving image by displaying a series of images or frames in succession.

[0929] "Distributing" refers to transmitting digital content over a network and making the content viewable on the receiving end.

[0930] "Description text of a dish" is a text about the contents, characteristics, ingredients, cooking method, etc. of a dish.

[0931] "Sign language gesture data" is a digital representation of sign language movements, and is typically data that indicates the position, direction, and movement of the hands and fingers.

[0932] "Visually impaired people" refers to people who have impaired vision and who have difficulty or are unable to understand visual information.

[0933] "Hearing impaired people" refers to people who have hearing impairments and who have difficulty or are unable to understand auditory information.

[0934] An "AI model" refers to an algorithm or computational model built using machine learning or deep learning, which performs data analysis and predictions.

[0935] "Via a network" refers to sending and receiving data via a communications infrastructure such as the Internet or a local area network.

[0936] This invention provides a system that enables visually and hearing impaired users to understand the details of dishes in food delivery services. Specifically, we apply technology that converts music lyrics into sign language to build a system that converts text describing dishes into sign language and generates and distributes videos using sign language by an avatar.

[0937] Basic system configuration

[0938] The server has the following functions:

[0939] 1. Text analysis:

[0940] The system analyzes the text description of a dish entered by the user and converts it into sign language gesture data using a generative AI model (using PyTorch and TensorFlow).

[0941] 2. Avatar movement generation:

[0942] Based on the analyzed sign language gesture data, a movement sequence for the avatar is generated, which specifically specifies the movements of each joint and body part required for sign language.

[0943] 3. Speech generation:

[0944] Use Google Cloud Speech-to-Text and Amazon Polly to convert the text of the cooking instructions into an audio file and add audio to the sign language video.

[0945] 4. Video Generation:

[0946] Using a video editing API such as OpenCV, sign language videos are generated so that the avatar's movements match the rhythm of the music.

[0947] 5. Delivery:

[0948] Using AWS S3 and CloudFront, the generated sign language videos are distributed to visually and hearing impaired people within a food delivery service application.

[0949] The user terminal has the following functions:

[0950] 1. Input interface:

[0951] It provides an interface for inputting text describing a dish. The user inputs the text and clicks the "Convert to sign language" button to start the process.

[0952] 2. Watching videos:

[0953] The distributed sign language video is played within the application, allowing the content to be conveyed to visually and hearing impaired people.

[0954] Processing flow

[0955] The user enters a text description of the dish: "This curry is made with special spices and is mild." The user's device sends the entered text data to the server, which converts the text into sign language gesture data. Next, the avatar movement generation function generates an avatar movement sequence based on the sign language gesture data. Furthermore, the audio file generated using Google Cloud Speech-to-Text and Amazon Polly is used to generate a sign language video of the avatar using a video editing API such as OpenCV. Finally, this video is distributed via AWS S3 and CloudFront and played on the user's device.

[0956] Prompt Sentence Examples

[0957] Below are some examples of dish descriptions:

[0958] "This curry is made with special spices and is not too spicy."

[0959] This pasta is made with plenty of seasonal vegetables.

[0960] This will enable users with visual or hearing impairments to understand details about dishes through food delivery services.

[0961] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0962] Step 1: User Input

[0963] The user inputs a text description of a dish through the food delivery service's application. For example, they input "This curry is made with special spices and is mild," and then click the "Sign Language Conversion" button. The input is sent to the device as text data.

[0964] Step 2: Send text data

[0965] The terminal converts the input text data into JSON format and sends it to the server. An appropriate communication protocol is used to ensure accurate data transfer. The input is text data and the output is JSON format data.

[0966] Step 3: Generate sign language gesture data

[0967] The server analyzes the received text data and inputs it into a generative AI model (using PyTorch or TensorFlow). The model converts the text into sign language gesture data. For example, it generates sign language gesture data for "curry." The input is text data, and the output is sign language gesture data.

[0968] Step 4: Generate avatar movement sequences

[0969] The server analyzes the sign language gesture data and generates a movement sequence for the avatar. A script is generated that specifically instructs the movement of each joint and body part, and the avatar's movement is determined based on this. The input is the sign language gesture data, and the output is the avatar's movement sequence.

[0970] Step 5: Speech generation

[0971] The server converts the text of the cooking instructions into an audio file using Google Cloud Speech-to-Text and Amazon Polly. The generated audio file is used to accompany the sign language video. The input is text data, and the output is an audio file.

[0972] Step 6: Video Generation

[0973] The server uses a video editing API such as OpenCV to combine the avatar's movement sequence based on the sign language gesture data with the generated audio file to generate a sign language video. This video is visually easy to understand, with the sign language and audio synchronized. The input is the avatar's movement sequence and audio file, and the output is the final sign language video.

[0974] Step 7: Publish your video

[0975] The server distributes the generated sign language video using AWS S3 and CloudFront, allowing users to watch the video on their devices. The input is the sign language video file, and the output is the video URL distributed over the network.

[0976] Step 8: Watch the video

[0977] The user device plays the sign language video delivered from the server and communicates the details of the dish to the visually and hearing impaired. The user can watch the video within the app and understand the explanation. The input is the video URL, and the output is the sign language video that is played.

[0978] This allows users with visual or hearing impairments to understand the details of the dishes and make appropriate choices.

[0979] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0980] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[0981] A specific embodiment of this system is described below.

[0982] First, the user inputs the lyrics of a song into the system. The user copies the lyrics from the music distribution service app and pastes them into the system's dedicated input form. The process begins by clicking the "Convert to Sign Language" button.

[0983] Next, the device sends the lyrics text data to the server, converts the input lyrics data into a JSON format request, and sends it to the server.

[0984] The server analyzes the received lyrics text and inputs it into the AI ​​model, which then analyzes the lyrics text and generates corresponding sign language gesture data, which is then used to generate avatar movements.

[0985] The server generates an avatar movement sequence based on this sign language gesture data, which specifically instructs the movement of each joint and body part of the avatar.The server then analyzes the music file and combines a timeline with the rhythm to synchronize the avatar movement sequence with the rhythm and tone of the music.

[0986] Furthermore, the server uses an emotion engine to recognize the user's emotions. The emotion engine uses facial recognition and voice analysis to detect the user's emotional state in real time. The recognized emotion information is reflected in the avatar's movements and sign language dance.

[0987] For example, if the user is expressing a happy emotion, the avatar's motion sequence may include more lively and joyful movements, whereas if the user is sad, the avatar may be controlled to include corresponding emotional movements.

[0988] The server then generates a dance video of the avatar based on the movement sequence. The video is visually easy to understand, with the avatar's movements and music perfectly matched. The video is saved in a common video format, such as MP4.

[0989] The server then distributes the generated video over the network, uploading the generated video file to a social networking platform, such as a video content sharing application.

[0990] Finally, users can watch the video on social media platforms. The video is presented in a short format, making it easy to watch and share. The video also includes relevant metadata (e.g., original lyrics, sign language meanings, etc.), making it easy for viewers to understand the sign language content.

[0991] For example, if a user inputs the lyrics "Spring, Come" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on the data, generates a dance video of the avatar in time with the music rhythm, and finally distributes it to a social media platform. During this process, the emotion engine recognizes the user's emotions in real time and adds corresponding emotional expressions to the avatar's movements.

[0992] The implementation of this system will enable sign language learning and provide a musical experience for the hearing-blind, and by adding actions that match the user's emotions, it is expected to further increase interest in sign language.

[0993] The processing flow will be explained below.

[0994] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[0995] The process flow will be explained in detail below for each step.

[0996] Step 1:

[0997] The user inputs the lyrics of a song into the system. For example, the user can copy the lyrics from a music streaming service app and paste them into the system's dedicated input form. This step begins by clicking the "Convert to Sign Language" button.

[0998] Step 2:

[0999] The device sends lyric text data to the server, converts the input lyric data into a JSON-formatted request, and sends it to the server. This communication is carried out via a secure protocol such as HTTPS.

[1000] Step 3:

[1001] The server receives the lyrics text and inputs it into the AI ​​model. The server analyzes the lyrics text and issues a request to the AI ​​model's API. The AI ​​model analyzes the lyrics and generates corresponding sign language gesture data.

[1002] Step 4:

[1003] The server receives the generated sign language data and stores it in an internal database, which is used later to generate avatar movements.

[1004] Step 5:

[1005] The server generates an avatar movement sequence from the sign language data.The server generates a movement sequence that specifically defines the joint and body movements of the avatar based on the sign language gesture data.

[1006] Step 6:

[1007] The server recognizes the user's emotions using an emotion engine. The emotion engine uses facial recognition and voice analysis of the user to detect the user's emotional state in real time. The emotion engine analyzes data input from the user's camera and microphone to identify whether the user is happy, sad, or other emotions.

[1008] Step 7:

[1009] The server adds emotional expressions to the avatar's motion sequence based on the recognized emotion information. For example, if the user is expressing a sense of enjoyment, the server adds more lively and joyful gestures to the avatar's motion. Conversely, if the user is sad, the server includes corresponding sadness-expressing motions.

[1010] Step 8:

[1011] The server generates a dance video of the avatar based on the sign language sequence, synchronized to the rhythm and tone of the music. The server analyzes the music file and calculates the timing based on the rhythm and tone. The server then matches the sign language sequence to the avatar's movements and generates the animation using 3D modeling software.

[1012] Step 9:

[1013] The server stores the generated dance video data and prepares it for distribution. The server saves the generated dance video in a format such as MP4. The server then converts the video file into a format suitable for distribution on social media and adds metadata (e.g., original lyrics, sign language explanations).

[1014] Step 10:

[1015] The server uploads the created dance video to the SNS platform. The server then sends the saved video file to the SNS platform's API for uploading.

[1016] Step 11:

[1017] Allow users to watch videos on the social media platform, which then makes the uploaded videos public and available for users to watch.

[1018] Step 12:

[1019] To make it easier for users to watch, the videos are distributed in short video format and related metadata (e.g., original lyrics, meaning of sign language) is also provided. The server automatically adds the video file metadata to the video description section of the social media platform.

[1020] For example, if a user inputs the lyrics "Spring, Come" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on the data, generates a dance video of the avatar in time with the music rhythm, and finally distributes it to a social media platform. During this process, the emotion engine recognizes the user's emotions in real time and adds corresponding emotional expressions to the avatar's movements.

[1021] The implementation of this system will enable sign language learning and provide a musical experience for the hearing-blind, and by adding actions that match the user's emotions, it is expected to further increase interest in sign language.

[1022] Example 2

[1023] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1024] Conventional systems that convert music lyrics into sign language simply convert lyrics into sign language and assign actions to an avatar, but do not take into account the user's emotions. As a result, the avatar's movements are monotonous and lack emotional appeal, and they are unable to sufficiently move or empathize with the viewer. There is also a need to provide a richer music experience for people with hearing and visual impairments.

[1025] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1026] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for synchronizing the generated avatar movements with the rhythm of the music to generate video, means for distributing the generated video via a network, and means for recognizing a user's emotions and adding emotional expressions to the avatar movements. This allows the avatar to be given movements and expressions that match the user's emotions, making it possible to provide video content that is more moving and resonates with viewers.

[1027] "Means for converting music lyrics into sign language" refers to a system or program for analyzing music lyrics text and converting it into sign language gesture data.

[1028] The "means for generating avatar movements based on sign language" refers to a script or algorithm for analyzing the generated sign language gesture data and generating a movement sequence for the avatar based thereon.

[1029] "Means for generating a video by matching the generated avatar's movements to the rhythm of music" refers to software or tools for adjusting the avatar's movement sequence to the rhythm or tone of music and generating it as a video.

[1030] "Means for distributing the generated video via a network" refers to a system or platform for distributing the generated video content via the Internet.

[1031] "Means for recognizing user emotions and adding emotional expressions to avatar movements" refers to an engine or algorithm that analyzes user emotions in real time and adds emotional elements to avatar movements based on that analysis.

[1032] A "generative AI model" is a model that uses machine learning and artificial intelligence techniques to convert input data (e.g., lyric text) into another format (e.g., sign language gesture data).

[1033] A "prompt" refers to a sentence or piece of data used as input for an AI model, and often includes specific instructions or questions.

[1034] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[1035] First, the user enters the lyrics of the music into the system. They copy the lyrics from the music distribution service app and paste them into the system's dedicated input form. The process begins by clicking the "Sign Language Conversion" button. The device used by the user is typically a smartphone or PC, and they access the system through a web browser.

[1036] Next, the device sends the lyric text data to the server. The device converts the input lyric data into a JSON format request and sends it to the server. The data is sent to a cloud server (e.g., Azure, AWS) using the HTTP protocol.

[1037] The server analyzes the received lyrics text and inputs it into a generative AI model. The software used includes natural language processing libraries (e.g., spaCy, NLTK) and machine learning frameworks (e.g., TensorFlow, PyTorch). The generative AI model analyzes the lyrics text and generates corresponding sign language gesture data. The generated sign language gesture data contains information that specifically instructs the movement of each joint and body part of the avatar.

[1038] The server generates an avatar's movement sequence based on this sign language gesture data. 3D content creation tools such as Blender and Unity are used to generate the 3D avatar's movements. A music analysis library (e.g., libROSA) is also used to analyze the music file and combine the movement sequence into a timeline in accordance with the rhythm and tone of the music.

[1039] Furthermore, the server uses an emotion engine to recognize the user's emotions. The emotion engine uses facial recognition technology (e.g., OpenCV, dlib) and voice analysis libraries (e.g., Google Cloud Speech-to-Text API) to detect the user's emotional state in real time. The recognized emotional information is reflected in the avatar's movements and sign language dance. For example, if the user is showing signs of enjoyment, a movement expressing joy will be added to the avatar's movement sequence.

[1040] The server generates a dance video of the avatar based on the movement sequence. It uses Blender or Unity to render the avatar's movements, and then uses FFmpeg to convert and save them as standard video files such as MP4.

[1041] The server then distributes the generated video over the network, and uploads the generated video file to social media platforms such as YouTube and Instagram, making it easily accessible to users and viewers.

[1042] Finally, users can watch the videos on social media platforms. The videos are presented in a short format, making them easy to watch and share. The videos also include relevant metadata, such as the original lyrics and sign language meanings, making it easy for viewers to understand the sign language content.

[1043] Here is a specific example: When a user inputs the lyrics "Haru yo, Koi" (Come Spring) and converts it into sign language, the process proceeds as follows:

[1044] Example prompt:

[1045] 1. The user accesses the system using a web browser, pastes the lyrics of "Haru yo, Koi" into the input form, and clicks the "Convert to Sign Language" button.

[1046] 2. The device converts the entered lyrics into a JSON format request and sends it to the server using the HTTP protocol.

[1047] 3. The server analyzes the lyrics text and generates sign language gesture data using a generative AI model.

[1048] 4. The server generates a movement sequence for the avatar based on the sign language gesture data, and the emotion engine recognizes the user's emotions and reflects them in the movement.

[1049] 5. The server generates the final avatar dance video and uploads it to a social media platform.

[1050] 6. The user watches the video on a social media platform and checks the meaning of the sign language.

[1051] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1052] Step 1:

[1053] The user enters the lyrics text.

[1054] Users access the system through a web browser, paste the music lyrics into the input form, and then click the "Convert to Sign Language" button.

[1055] Input: Lyric text (e.g. "Haru yo, Koi" (Come Spring))

[1056] Output: Signals the start of a request to the system

[1057] Step 2:

[1058] The terminal transmits the lyrics text data to the server.

[1059] The device receives the lyrics data entered by the user, converts it into a JSON-formatted request, and then sends the request to the server using the HTTP protocol.

[1060] Input: Lyric text

[1061] Data processing: Convert lyrics data into JSON format

[1062] Output: HTTP request to the server

[1063] Step 3:

[1064] The server parses the lyrics text.

[1065] The server parses the received JSON request and extracts the lyrics text, then uses a natural language processing library (e.g., spaCy) to analyze the lyrics and extract important keywords and phrases.

[1066] Input: Lyrics data in JSON format

[1067] Data processing: Lyric text analysis, keyword extraction

[1068] Output: Analyzed text data (keywords and phrases)

[1069] Step 4:

[1070] The server generates sign language gesture data.

[1071] The server uses a generative AI model (e.g., a TensorFlow model) to convert the parsed text data into sign language gesture data.

[1072] Input: Parsed text data

[1073] Data Computation: Transformation with Generative AI Models

[1074] Output: Sign language gesture data

[1075] Step 5:

[1076] The server generates the avatar movement sequence.

[1077] The server generates avatar movement sequences based on sign language gesture data using 3D content creation tools such as Blender and Unity, and analyzes music files to combine movement sequences into a timeline in sync with the rhythm and tone of the music.

[1078] Input: Sign language gesture data, music files

[1079] Data processing: generating motion sequences and combining timelines

[1080] Output: Avatar movement sequence

[1081] Step 6:

[1082] The server uses an emotion engine to recognize the user's emotion.

[1083] The server uses facial recognition technology (e.g., OpenCV, dlib) and voice analysis libraries (e.g., Google Cloud Speech-to-Text API) to detect the user's emotional state in real time. The recognized emotional information is reflected in the avatar's behavior.

[1084] Input: User's face image, voice data

[1085] Data Computing: Applying Emotion Recognition Algorithms

[1086] Output: User's emotional state information

[1087] Step 7:

[1088] The server generates a dance video of the avatar.

[1089] The server generates a dance video of the avatar based on the final movement sequence. It renders the avatar's movements using Blender or Unity and exports them to a video file (e.g., MP4 format) using FFmpeg.

[1090] Input: Action sequence, emotional state information

[1091] Data processing: video rendering, file export

[1092] Output: Dance video file

[1093] Step 8:

[1094] The server distributes the generated video over the network.

[1095] The server uploads the generated video file to a social media platform (e.g. YouTube, Instagram), allowing users to access the video file.

[1096] Input: Dance video file

[1097] Data transmission: Video upload

[1098] Output: Public videos on social media platforms

[1099] Step 9:

[1100] A user watches a video on a social media platform.

[1101] Users access the social media platform and watch videos of the generated avatar dancing, with associated metadata such as the original lyrics and sign language meanings displayed alongside the video.

[1102] Input: Access to social media platforms

[1103] Output: Watchable dance video and associated metadata

[1104] (Application example 2)

[1105] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1106] Conventional sign language translation systems cannot reflect the user's emotions when converting music lyrics into sign language and expressing them with an avatar, making it difficult to add emotional expression to the sign language movements or avatar dance. Furthermore, even if the generated video content is visually rich, it lacks movements that correspond to the user's emotions, limiting the viewing experience.

[1107] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1108] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for generating a video by synchronizing the generated avatar movements with the rhythm of the music, means for recognizing the user's emotions and adding emotional expressions to the avatar movements, and means for distributing the generated video via a network. This makes it possible to generate a sign language dance video that incorporates expressions according to the user's emotions, providing a richer viewing experience.

[1109] "Music" is an art form that expresses human emotion and experience in the form of sound, and includes rhythm, melody, and harmony.

[1110] "Lyrics" are a part of music that expresses the content and theme of a song in words.

[1111] "Sign language" is a visual, gestural language used for communication between people with hearing impairments, using hand and finger movements and facial expressions.

[1112] An "avatar" refers to a virtual person or character that symbolically represents a user or character in a digital space.

[1113] "Movement" refers to the physical movements and gestures performed by the avatar, including things like sign language and dancing.

[1114] "Video" is a media format that visually recreates movement by displaying a series of still images in succession.

[1115] "Rhythm" refers to the pattern of changes in the length and dynamics of notes in music, and is an element that gives a piece of music a sense of unity and movement.

[1116] "Emotions" refer to the psychological states that humans feel in response to specific situations or events, and include joy, sadness, anger, etc.

[1117] A "network" is a system of computers and communication devices connected together to share information and data.

[1118] "Distribution" refers to the act of providing or transmitting digital content to users over a network.

[1119] A "server" is a computer system that provides data and services over a network.

[1120] A "generative AI model" is a model generated by artificial intelligence and trained to predict or generate appropriate outputs for specific inputs.

[1121] A "prompt sentence" is an input sentence provided to a generative AI model to obtain a desired output.

[1122] This system converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates a function to recognize the user's emotions and add emotional expressions to the avatar's movements.

[1123] System Configuration

[1124] The system consists of the following main components: a user device, a server, a generative AI model, an emotion recognition engine, and a video distribution platform.

[1125] Program processing description

[1126] 1. User Device

[1127] - Lyrics text input: The user device provides an input form for entering music lyrics. The user enters the lyrics text of their favorite song into this form and clicks the "Convert to sign language" button.

[1128] - Emotion recognition: The user's face is captured in real time through the camera on the user's device, and the emotion recognition engine analyzes the user's emotions.

[1129] 2. Data Transmission

[1130] - The device converts the entered lyrics data into JSON format and sends it to the server. At the same time, the user's emotional data is also sent.

[1131] 3. Server Processing

[1132] - The server analyzes the received lyrics text and converts them into sign language gesture data using a generative AI model (e.g., OpenAI's GPT-4). This conversion process uses prompts.

[1133] - Use an emotion recognition engine (e.g., Microsoft Azure Emotion Recognition API) to obtain the user's emotional information and reflect it in the avatar's behavior.

[1134] 4. Avatar Motion Generation

[1135] The server generates a movement sequence for the avatar based on the generated sign language gesture data, which specifically instructs the movements of each joint and body part of the avatar.

[1136] - Perform rhythmic analysis of the music (e.g., librosa library) and synchronize the avatar's movement sequences to the music.

[1137] 5. Video Generation

[1138] - The server generates a dance video by combining the avatar's movement sequence with music, using a video generation tool (e.g., OpenCV and ffmpeg library). The generated video is saved in MP4 format.

[1139] 6. Video streaming

[1140] - The server distributes the generated videos over the network, and uploads them to social media platforms (e.g. YouTube API, Facebook API) so that users can easily watch and share them within the app.

[1141] Examples of concrete examples and prompts

[1142] As a concrete example, when the lyrics "Haru yo, Koi" (Come Spring) are entered by a user, the generative AI model generates sign language gesture data using the following prompt sentence:

[1143] Input lyrics:

[1144] "Come Spring"

[1145] Prompt statement:

[1146] Convert the following lyrics into sign language and generate sign language gesture data for the avatar. Also, please consider that the user's emotion is "joy." Output the generated gesture data in JSON format.

[1147] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1148] Step 1:

[1149] The user inputs music lyric text into the terminal and clicks the "Convert to Sign Language" button.

[1150] Input: Music lyric text

[1151] Output: Lyric text input confirmation

[1152] First, the user enters the lyrics of the music into the application's dedicated input form and clicks the "Convert to Sign Language" button. At this time, the user's camera is activated and captures facial images in real time, which prepares the emotion recognition engine to analyze the user's emotions.

[1153] Step 2:

[1154] The device converts the lyrics text into JSON format and sends it to the server, along with the user's facial image data.

[1155] Input: Lyric text, user's facial image data

[1156] Output: Sending lyrics data in JSON format and facial video data to the server

[1157] The device converts the entered lyrics text into JSON format and sends it to a cloud server. A video of the user's face is also transferred at the same time, allowing subsequent processing to proceed smoothly.

[1158] Step 3:

[1159] The server analyzes the lyric text and generates sign language gesture data using a generative AI model.

[1160] Input: Lyrics data in JSON format

[1161] Output: Sign language gesture data

[1162] The server parses the received JSON-formatted lyrics data and converts the lyrics text into sign language gesture data using a generative AI model (e.g., OpenAI's GPT-4). During this process, an appropriate prompt is passed to the generative AI model.

[1163] Step 4:

[1164] The server uses an emotion recognition engine to analyze the user's emotions and obtain emotion data.

[1165] Input: Facial image data

[1166] Output: Emotion data

[1167] The server uses an emotion recognition engine (e.g., Microsoft Azure Emotion Recognition API) to analyze the facial video data and identify the user's emotions. The emotion data is saved so that it can be reflected in the avatar's behavior.

[1168] Step 5:

[1169] The server generates an avatar's movement sequence based on the sign language gesture data and emotion data.

[1170] Input: Sign language gesture data, emotion data

[1171] Output: Avatar movement sequence

[1172] The server combines the sign language gesture data and emotion data to generate a movement sequence for the avatar, which generates realistic and emotional avatar movements based on the user's emotions. Specific instructions for the movements of each joint and body part are created.

[1173] Step 6:

[1174] The server analyzes the rhythm of the music and synchronizes the avatar's movement sequence with the music.

[1175] Input: Action sequences, music files

[1176] Output: rhythm-synchronized movement sequence

[1177] The server analyzes the music files to identify the rhythm (e.g., using the librosa library), then combines the rhythmic timeline with a movement sequence to create engaging, music-synchronized avatar movements.

[1178] Step 7:

[1179] The server combines the avatar's movement sequence with music to generate a video.

[1180] Input: rhythm-synchronized movement sequences, music files

[1181] Output: Generated video file

[1182] The server finally generates a dance video by combining the avatar's movement sequence with music. This video generation uses libraries such as OpenCV and ffmpeg to generate a video file in MP4 format.

[1183] Step 8:

[1184] The server distributes the generated video to the social media platform via the network.

[1185] Input: Generated video file

[1186] Output: Upload to social media platforms

[1187] The server uploads the generated video to a social networking platform (e.g., YouTube API, Facebook API) for users to easily access, watch, and share, allowing users and viewers to enjoy the emotional avatar movements and music.

[1188] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1189] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1190] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1191] [Fourth embodiment]

[1192] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1193] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1194] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1195] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1196] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1197] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1198] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1199] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1200] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1201] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1202] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1203] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1204] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1205] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing in accompaniment with the sign language. This system performs a series of processes, including inputting lyrics text, converting them into sign language, generating avatar movements, generating videos, and distributing them.

[1206] First, the user inputs the lyrics of the music into the system using their own device. For example, the user can copy the lyrics from a music distribution service app and paste them into the system's dedicated input form. Then, the process begins by clicking the "Sign Language Conversion" button.

[1207] Next, the device sends the lyrics text data to the server, where it converts the input lyrics data into JSON format and sends it to the server using a dedicated communication protocol for fast and reliable data transfer.

[1208] The server analyzes the received lyrics text and inputs it into a pre-trained AI model, which then analyzes the lyrics text and generates corresponding sign language gesture data, which is then used to generate avatar movements.

[1209] The server generates an avatar movement sequence based on this sign language gesture data, which specifically instructs the movement of each joint and body part of the avatar.The server then analyzes the music file to synchronize the avatar movement sequence with the rhythm and tone of the music, and combines a timeline with the rhythm.

[1210] The server then generates a dance video of the avatar based on the movement sequence, matching the avatar's movements with the music to create a visually appealing video. The video is saved in a common video format, such as MP4.

[1211] The server then distributes the generated video over the network, uploading the video file to a dedicated social networking platform, such as a video content sharing application.

[1212] Finally, users can watch the video on social media platforms. The video is presented in a short format, making it easy to watch and share. The video also includes relevant metadata (e.g., original lyrics, sign language meanings, etc.), making it easy for viewers to understand the sign language content.

[1213] As a concrete example, if a user inputs the lyrics "Come Spring" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on that, generates a dance video of the avatar in time with the rhythm of the music, and finally distributes it to a social media platform.

[1214] The implementation of this system will enable people with visual and hearing impairments to learn sign language and to experience music, and is expected to increase interest in sign language.

[1215] The processing flow will be explained below.

[1216] Step 1:

[1217] The user inputs the lyrics of a song into the system. The user copies the lyrics from a music distribution service app and pastes them into the system's dedicated input form. Then, the user clicks the "Convert to Sign Language" button.

[1218] Step 2:

[1219] The device sends lyric text data to the server. The device converts the input lyric data into a JSON format request and sends it to the server.

[1220] Step 3:

[1221] The server receives the lyrics text and inputs it into the AI ​​model. The server analyzes the lyrics text and issues a request to the AI ​​model's API. The AI ​​model analyzes the lyrics and generates corresponding sign language gesture data.

[1222] Step 4:

[1223] The server receives the generated sign language data and stores it in an internal database, which is used to generate movement sequences for the avatar.

[1224] Step 5:

[1225] The server generates an avatar's movement sequence from the sign language data. The server executes a script that generates a sequence corresponding to the avatar's joint and body movements based on the sign language gesture data.

[1226] Step 6:

[1227] The server generates a dance video of the avatar based on the sign language sequence and synchronizes it with the rhythm and tone of the music. The server analyzes the music file and calculates the timing based on the rhythm and tone. The server then reflects the sign language sequence in the avatar's movements according to the timing and generates the animation using 3D modeling software.

[1228] Step 7:

[1229] The server saves the generated dance video data and prepares it for distribution. The server saves the generated dance video in a format such as MP4. The server converts the video file into a format suitable for distribution on social media and adds metadata (e.g., original lyrics, sign language explanations).

[1230] Step 8:

[1231] The server uploads the created dance video to the SNS platform. The server then sends the saved video file to the SNS platform's API for uploading.

[1232] Step 9:

[1233] Allow users to watch videos on the social media platform, which then makes the uploaded videos public and available for users to watch.

[1234] Step 10:

[1235] To make it easier for users to watch, the videos are distributed in short video format and related metadata (e.g., original lyrics, meaning of sign language) is also provided. The server automatically adds the video file metadata to the video description section of the social media platform.

[1236] Example 1

[1237] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1238] With conventional technology, it has been difficult to convert music lyrics into sign language and efficiently generate and distribute videos of avatars dancing using sign language. Furthermore, there have been limited ways for the hearing-blind to enjoy music, and there has also been a lack of methods to promote learning sign language. This has limited musical experiences and learning opportunities for them.

[1239] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1240] In this invention, the server includes means for converting lyric text data into JSON format and sending it to the server, means for analyzing the lyric text data using an AI model and converting it into sign language gesture data, and means for generating an avatar movement sequence based on the sign language gesture data. This allows the generation of an avatar movement sequence in sync with the rhythm of music, providing a means for the hearing-blind to enjoy music and promoting the learning of sign language.

[1241] "Lyric text data" refers to data in which music lyrics are saved as text information.

[1242] "Sign language gesture data" is data that expresses each gesture or movement in sign language.

[1243] "JSON format" is an abbreviation for JavaScript Object Notation, and is a lightweight data exchange format that consists of data in key-value pairs.

[1244] An "AI model" is an algorithm or system trained using artificial intelligence techniques to perform a specific task or analysis.

[1245] An "avatar" refers to a virtual person or character in a digital environment.

[1246] A "movement sequence" is a series of commands that specify the movement of each joint and body part of the avatar along a time axis.

[1247] "3D modeling software" is software for creating, editing, and displaying three-dimensional models.

[1248] The "MP4 format" is a common digital file format for storing video and audio.

[1249] A "network" is a system of computers and other devices connected to exchange data with each other.

[1250] An "SNS platform" is an online service that allows users to post, share, and view content.

[1251] "Sign language" is a system of visual gestures and movements that enables deaf people to communicate.

[1252] The present invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing in accompaniment with the sign language. An embodiment of this system will be described in detail below.

[1253] First, the user inputs the lyrics of the music using their own device. The user copies the lyrics from a music distribution service application or similar and pastes them into the designated input form. Then, the process begins by clicking the "Sign Language Conversion" button.

[1254] Next, the device converts the entered lyrics text data into JSON format and sends it to the server. The device uses a dedicated communication protocol for this data conversion and transmission.

[1255] The server analyzes the received lyric text data using a pre-trained generative AI model, which converts the lyric text into sign language gesture data, which is then used to generate a movement sequence for the avatar.

[1256] The server then generates a motion sequence for the avatar based on the sign language gesture data, which specifically instructs the movements of each joint and body part of the avatar. It also analyzes the music file and combines a timeline with the rhythm of the music to synchronize the avatar's movements with the rhythm of the music.

[1257] The server then generates a dance video of the avatar based on the movement sequence, using 3D modeling software. The video, which matches the avatar's movements with the music, is visually appealing and saved in a common video format such as MP4.

[1258] Finally, the server distributes the generated video over the network. The video file is uploaded to a dedicated social networking platform, where users can watch the video. The video also includes relevant metadata, such as the original lyrics and sign language meanings, making it easy for viewers to understand the sign language content.

[1259] For example, if a user inputs the lyrics "Spring Comes" into the system and clicks the "Sign Language Conversion" button, the server converts the lyrics into sign language and generates sign language gesture data. Next, based on the generated sign language gesture data, the server creates a movement sequence for the avatar and generates a video of the avatar dancing to the rhythm of the music. Finally, this video is distributed to a social networking platform.

[1260] Example prompt sentence:

[1261] "Enter the following lyrics, 'Spring Comes,' into the system, convert them into sign language, generate a dance video with an avatar, and post it on social media."

[1262] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1263] Step 1:

[1264] A user uses his / her terminal to input the lyrics text of the music.

[1265] Specific operation: The user copies lyrics from a music distribution service application, pastes them into a dedicated input form, and clicks the "Convert to sign language" button.

[1266] Input: Music lyric text

[1267] Output: Trigger the start of processing by clicking the "Sign Language Conversion" button

[1268] Step 2:

[1269] The device converts the entered lyrics text data into JSON format and sends it to the server.

[1270] Specific operation: The terminal converts the text data into JSON format and sends it to the server using a dedicated communication protocol.

[1271] Input: Lyric text entered by the user

[1272] Output: Lyrics text data in JSON format sent to the server

[1273] Step 3:

[1274] The server analyzes the received lyrics text data.

[1275] Specific operation: The server analyzes the received JSON format lyrics text data using the analysis module.

[1276] Input: Lyrics text data in JSON format

[1277] Output: Parsed lyrics text data

[1278] Step 4:

[1279] The server inputs the analyzed lyrics text data into a generative AI model to generate sign language gesture data.

[1280] Specific operation: Text data is input into the generative AI model, and sign language gesture data is generated.

[1281] Input: Parsed lyrics text data

[1282] Output: Sign language gesture data

[1283] Step 5:

[1284] The server generates a movement sequence for the avatar based on the sign language gesture data.

[1285] Specific movements: Based on sign language gesture data, a movement sequence is generated that specifically instructs the movement of each joint and body part of the avatar.

[1286] Input: Sign language gesture data

[1287] Output: Avatar movement sequence

[1288] Step 6:

[1289] The server analyzes the music files and combines the movement sequences with rhythm.

[1290] Specific Actions: Analyze music files and combine action sequences into a timeline based on the rhythm of the music.

[1291] Input: Avatar movement sequence, music file

[1292] Output: Motion sequence synchronized with the music rhythm

[1293] Step 7:

[1294] The server generates a dance video based on the movement sequence.

[1295] Specific Movement: Using 3D modeling software, a dance video is generated that combines an avatar's movement sequence with music.

[1296] Input: A sequence of movements synchronized with the musical rhythm

[1297] Output: Generated dance video (MP4 format, etc.)

[1298] Step 8:

[1299] The server distributes the generated video over the network.

[1300] Specific operation: Upload the generated video file to a social media platform.

[1301] Input: Generated dance video (e.g. MP4 format)

[1302] Output: Upload completion notification to SNS platform

[1303] Step 9:

[1304] Users watch videos on social media platforms.

[1305] Specific Action: A user visits a social media platform to watch and share a video.

[1306] Input: Video link on social media platform

[1307] Output: Videos watched by users, videos shared

[1308] (Application example 1)

[1309] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1310] Conventional food delivery services have the problem that it is difficult for visually or hearing impaired users to understand the details of the dishes. In addition, there are many cases where explanations are not available in sign language, and a sign language interpreter is required. There is a need for a means to enable visually or audibly impaired users to understand detailed descriptions of dishes.

[1311] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1312] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for generating a video in which the generated avatar movements are synchronized with the rhythm of the music, means for distributing the generated video via a network, means for converting text describing dishes into sign language and generating a video in which an avatar uses sign language, and means for distributing the generated sign language video to visually impaired users, thereby enabling visually or hearing impaired users to understand the details of dishes when using a food delivery service.

[1313] "Musical lyrics" are text, usually consisting of words, that is sung to accompany a melody in a musical performance or song.

[1314] "Sign language" is a visual language that primarily communicates through hand and finger movements and facial expressions.

[1315] An "avatar" is a character or graphical representation that represents a user and acts on behalf of the user in a virtual space or on a computer.

[1316] "Generating behavior" refers to creating a series of actions or poses for an avatar or character on a computer.

[1317] "Musical rhythm" refers to the temporal pattern or tempo of music.

[1318] "Generating video" refers to creating a moving image by displaying a series of images or frames in succession.

[1319] "Distributing" refers to transmitting digital content over a network and making the content viewable on the receiving end.

[1320] "Description text of a dish" is a text about the contents, characteristics, ingredients, cooking method, etc. of a dish.

[1321] "Sign language gesture data" is a digital representation of sign language movements, and is typically data that indicates the position, direction, and movement of the hands and fingers.

[1322] "Visually impaired people" refers to people who have impaired vision and who have difficulty or are unable to understand visual information.

[1323] "Hearing impaired people" refers to people who have hearing impairments and who have difficulty or are unable to understand auditory information.

[1324] An "AI model" refers to an algorithm or computational model built using machine learning or deep learning, which performs data analysis and predictions.

[1325] "Via a network" refers to sending and receiving data via a communications infrastructure such as the Internet or a local area network.

[1326] This invention provides a system that enables visually and hearing impaired users to understand the details of dishes in food delivery services. Specifically, we apply technology that converts music lyrics into sign language to build a system that converts text describing dishes into sign language and generates and distributes videos using sign language by an avatar.

[1327] Basic system configuration

[1328] The server has the following functions:

[1329] 1. Text analysis:

[1330] The system analyzes the text description of a dish entered by the user and converts it into sign language gesture data using a generative AI model (using PyTorch and TensorFlow).

[1331] 2. Avatar movement generation:

[1332] Based on the analyzed sign language gesture data, a movement sequence for the avatar is generated, which specifically specifies the movements of each joint and body part required for sign language.

[1333] 3. Speech generation:

[1334] Use Google Cloud Speech-to-Text and Amazon Polly to convert the text of the cooking instructions into an audio file and add audio to the sign language video.

[1335] 4. Video Generation:

[1336] Using a video editing API such as OpenCV, sign language videos are generated so that the avatar's movements match the rhythm of the music.

[1337] 5. Delivery:

[1338] Using AWS S3 and CloudFront, the generated sign language videos are distributed to visually and hearing impaired people within a food delivery service application.

[1339] The user terminal has the following functions:

[1340] 1. Input interface:

[1341] It provides an interface for inputting text describing a dish. The user inputs the text and clicks the "Convert to sign language" button to start the process.

[1342] 2. Watching videos:

[1343] The distributed sign language video is played within the application, allowing the content to be conveyed to visually and hearing impaired people.

[1344] Processing flow

[1345] The user enters a text description of the dish: "This curry is made with special spices and is mild." The user's device sends the entered text data to the server, which converts the text into sign language gesture data. Next, the avatar movement generation function generates an avatar movement sequence based on the sign language gesture data. Furthermore, the audio file generated using Google Cloud Speech-to-Text and Amazon Polly is used to generate a sign language video of the avatar using a video editing API such as OpenCV. Finally, this video is distributed via AWS S3 and CloudFront and played on the user's device.

[1346] Prompt Sentence Examples

[1347] Below are some examples of dish descriptions:

[1348] "This curry is made with special spices and is not too spicy."

[1349] This pasta is made with plenty of seasonal vegetables.

[1350] This will enable users with visual or hearing impairments to understand details about dishes through food delivery services.

[1351] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1352] Step 1: User Input

[1353] The user inputs a text description of a dish through the food delivery service's application. For example, they input "This curry is made with special spices and is mild," and then click the "Sign Language Conversion" button. The input is sent to the device as text data.

[1354] Step 2: Send text data

[1355] The terminal converts the input text data into JSON format and sends it to the server. An appropriate communication protocol is used to ensure accurate data transfer. The input is text data and the output is JSON format data.

[1356] Step 3: Generate sign language gesture data

[1357] The server analyzes the received text data and inputs it into a generative AI model (using PyTorch or TensorFlow). The model converts the text into sign language gesture data. For example, it generates sign language gesture data for "curry." The input is text data, and the output is sign language gesture data.

[1358] Step 4: Generate avatar movement sequences

[1359] The server analyzes the sign language gesture data and generates a movement sequence for the avatar. A script is generated that specifically instructs the movement of each joint and body part, and the avatar's movement is determined based on this. The input is the sign language gesture data, and the output is the avatar's movement sequence.

[1360] Step 5: Speech generation

[1361] The server converts the text of the cooking instructions into an audio file using Google Cloud Speech-to-Text and Amazon Polly. The generated audio file is used to accompany the sign language video. The input is text data, and the output is an audio file.

[1362] Step 6: Video Generation

[1363] The server uses a video editing API such as OpenCV to combine the avatar's movement sequence based on the sign language gesture data with the generated audio file to generate a sign language video. This video is visually easy to understand, with the sign language and audio synchronized. The input is the avatar's movement sequence and audio file, and the output is the final sign language video.

[1364] Step 7: Publish your video

[1365] The server distributes the generated sign language video using AWS S3 and CloudFront, allowing users to watch the video on their devices. The input is the sign language video file, and the output is the video URL distributed over the network.

[1366] Step 8: Watch the video

[1367] The user device plays the sign language video delivered from the server and communicates the details of the dish to the visually and hearing impaired. The user can watch the video within the app and understand the explanation. The input is the video URL, and the output is the sign language video that is played.

[1368] This allows users with visual or hearing impairments to understand the details of the dishes and make appropriate choices.

[1369] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1370] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[1371] A specific embodiment of this system is described below.

[1372] First, the user inputs the lyrics of a song into the system. The user copies the lyrics from the music distribution service app and pastes them into the system's dedicated input form. The process begins by clicking the "Convert to Sign Language" button.

[1373] Next, the device sends the lyrics text data to the server, converts the input lyrics data into a JSON format request, and sends it to the server.

[1374] The server analyzes the received lyrics text and inputs it into the AI ​​model, which then analyzes the lyrics text and generates corresponding sign language gesture data, which is then used to generate avatar movements.

[1375] The server generates an avatar movement sequence based on this sign language gesture data, which specifically instructs the movement of each joint and body part of the avatar.The server then analyzes the music file and combines a timeline with the rhythm to synchronize the avatar movement sequence with the rhythm and tone of the music.

[1376] Furthermore, the server uses an emotion engine to recognize the user's emotions. The emotion engine uses facial recognition and voice analysis to detect the user's emotional state in real time. The recognized emotion information is reflected in the avatar's movements and sign language dance.

[1377] For example, if the user is expressing a happy emotion, the avatar's motion sequence may include more lively and joyful movements, whereas if the user is sad, the avatar may be controlled to include corresponding emotional movements.

[1378] The server then generates a dance video of the avatar based on the movement sequence. The video is visually easy to understand, with the avatar's movements and music perfectly matched. The video is saved in a common video format, such as MP4.

[1379] The server then distributes the generated video over the network, uploading the generated video file to a social networking platform, such as a video content sharing application.

[1380] Finally, users can watch the video on social media platforms. The video is presented in a short format, making it easy to watch and share. The video also includes relevant metadata (e.g., original lyrics, sign language meanings, etc.), making it easy for viewers to understand the sign language content.

[1381] For example, if a user inputs the lyrics "Spring, Come" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on the data, generates a dance video of the avatar in time with the music rhythm, and finally distributes it to a social media platform. During this process, the emotion engine recognizes the user's emotions in real time and adds corresponding emotional expressions to the avatar's movements.

[1382] The implementation of this system will enable sign language learning and provide a musical experience for the hearing-blind, and by adding actions that match the user's emotions, it is expected to further increase interest in sign language.

[1383] The processing flow will be explained below.

[1384] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[1385] The process flow will be explained in detail below for each step.

[1386] Step 1:

[1387] The user inputs the lyrics of a song into the system. For example, the user can copy the lyrics from a music streaming service app and paste them into the system's dedicated input form. This step begins by clicking the "Convert to Sign Language" button.

[1388] Step 2:

[1389] The device sends lyric text data to the server, converts the input lyric data into a JSON-formatted request, and sends it to the server. This communication is carried out via a secure protocol such as HTTPS.

[1390] Step 3:

[1391] The server receives the lyrics text and inputs it into the AI ​​model. The server analyzes the lyrics text and issues a request to the AI ​​model's API. The AI ​​model analyzes the lyrics and generates corresponding sign language gesture data.

[1392] Step 4:

[1393] The server receives the generated sign language data and stores it in an internal database, which is used later to generate avatar movements.

[1394] Step 5:

[1395] The server generates an avatar movement sequence from the sign language data.The server generates a movement sequence that specifically defines the joint and body movements of the avatar based on the sign language gesture data.

[1396] Step 6:

[1397] The server recognizes the user's emotions using an emotion engine. The emotion engine uses facial recognition and voice analysis of the user to detect the user's emotional state in real time. The emotion engine analyzes data input from the user's camera and microphone to identify whether the user is happy, sad, or other emotions.

[1398] Step 7:

[1399] The server adds emotional expressions to the avatar's motion sequence based on the recognized emotion information. For example, if the user is expressing a sense of enjoyment, the server adds more lively and joyful gestures to the avatar's motion. Conversely, if the user is sad, the server includes corresponding sadness-expressing motions.

[1400] Step 8:

[1401] The server generates a dance video of the avatar based on the sign language sequence, synchronized to the rhythm and tone of the music. The server analyzes the music file and calculates the timing based on the rhythm and tone. The server then matches the sign language sequence to the avatar's movements and generates the animation using 3D modeling software.

[1402] Step 9:

[1403] The server stores the generated dance video data and prepares it for distribution. The server saves the generated dance video in a format such as MP4. The server then converts the video file into a format suitable for distribution on social media and adds metadata (e.g., original lyrics, sign language explanations).

[1404] Step 10:

[1405] The server uploads the created dance video to the SNS platform. The server then sends the saved video file to the SNS platform's API for uploading.

[1406] Step 11:

[1407] Allow users to watch videos on the social media platform, which then makes the uploaded videos public and available for users to watch.

[1408] Step 12:

[1409] To make it easier for users to watch, the videos are distributed in short video format and related metadata (e.g., original lyrics, meaning of sign language) is also provided. The server automatically adds the video file metadata to the video description section of the social media platform.

[1410] For example, if a user inputs the lyrics "Spring, Come" and converts them into sign language, the server converts the lyrics into sign language, generates sign language gesture data, creates an avatar movement sequence based on the data, generates a dance video of the avatar in time with the music rhythm, and finally distributes it to a social media platform. During this process, the emotion engine recognizes the user's emotions in real time and adds corresponding emotional expressions to the avatar's movements.

[1411] The implementation of this system will enable sign language learning and provide a musical experience for the hearing-blind, and by adding actions that match the user's emotions, it is expected to further increase interest in sign language.

[1412] Example 2

[1413] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1414] Conventional systems that convert music lyrics into sign language simply convert lyrics into sign language and assign actions to an avatar, but do not take into account the user's emotions. As a result, the avatar's movements are monotonous and lack emotional appeal, and they are unable to sufficiently move or empathize with the viewer. There is also a need to provide a richer music experience for people with hearing and visual impairments.

[1415] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1416] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for synchronizing the generated avatar movements with the rhythm of the music to generate video, means for distributing the generated video via a network, and means for recognizing a user's emotions and adding emotional expressions to the avatar movements. This allows the avatar to be given movements and expressions that match the user's emotions, making it possible to provide video content that is more moving and resonates with viewers.

[1417] "Means for converting music lyrics into sign language" refers to a system or program for analyzing music lyrics text and converting it into sign language gesture data.

[1418] The "means for generating avatar movements based on sign language" refers to a script or algorithm for analyzing the generated sign language gesture data and generating a movement sequence for the avatar based thereon.

[1419] "Means for generating a video by matching the generated avatar's movements to the rhythm of music" refers to software or tools for adjusting the avatar's movement sequence to the rhythm or tone of music and generating it as a video.

[1420] "Means for distributing the generated video via a network" refers to a system or platform for distributing the generated video content via the Internet.

[1421] "Means for recognizing user emotions and adding emotional expressions to avatar movements" refers to an engine or algorithm that analyzes user emotions in real time and adds emotional elements to avatar movements based on that analysis.

[1422] A "generative AI model" is a model that uses machine learning and artificial intelligence techniques to convert input data (e.g., lyric text) into another format (e.g., sign language gesture data).

[1423] A "prompt" refers to a sentence or piece of data used as input for an AI model, and often includes specific instructions or questions.

[1424] This invention relates to a system that converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates an emotion engine that recognizes the user's emotions, and is characterized by adding emotional expressions to the avatar's movements based on the user's emotions.

[1425] First, the user enters the lyrics of the music into the system. They copy the lyrics from the music distribution service app and paste them into the system's dedicated input form. The process begins by clicking the "Sign Language Conversion" button. The device used by the user is typically a smartphone or PC, and they access the system through a web browser.

[1426] Next, the device sends the lyric text data to the server. The device converts the input lyric data into a JSON format request and sends it to the server. The data is sent to a cloud server (e.g., Azure, AWS) using the HTTP protocol.

[1427] The server analyzes the received lyrics text and inputs it into a generative AI model. The software used includes natural language processing libraries (e.g., spaCy, NLTK) and machine learning frameworks (e.g., TensorFlow, PyTorch). The generative AI model analyzes the lyrics text and generates corresponding sign language gesture data. The generated sign language gesture data contains information that specifically instructs the movement of each joint and body part of the avatar.

[1428] The server generates an avatar's movement sequence based on this sign language gesture data. 3D content creation tools such as Blender and Unity are used to generate the 3D avatar's movements. A music analysis library (e.g., libROSA) is also used to analyze the music file and combine the movement sequence into a timeline in accordance with the rhythm and tone of the music.

[1429] Furthermore, the server uses an emotion engine to recognize the user's emotions. The emotion engine uses facial recognition technology (e.g., OpenCV, dlib) and voice analysis libraries (e.g., Google Cloud Speech-to-Text API) to detect the user's emotional state in real time. The recognized emotional information is reflected in the avatar's movements and sign language dance. For example, if the user is showing signs of enjoyment, a movement expressing joy will be added to the avatar's movement sequence.

[1430] The server generates a dance video of the avatar based on the movement sequence. It uses Blender or Unity to render the avatar's movements, and then uses FFmpeg to convert and save them as standard video files such as MP4.

[1431] The server then distributes the generated video over the network, and uploads the generated video file to social media platforms such as YouTube and Instagram, making it easily accessible to users and viewers.

[1432] Finally, users can watch the videos on social media platforms. The videos are presented in a short format, making them easy to watch and share. The videos also include relevant metadata, such as the original lyrics and sign language meanings, making it easy for viewers to understand the sign language content.

[1433] Here is a specific example: When a user inputs the lyrics "Haru yo, Koi" (Come Spring) and converts it into sign language, the process proceeds as follows:

[1434] Example prompt:

[1435] 1. The user accesses the system using a web browser, pastes the lyrics of "Haru yo, Koi" into the input form, and clicks the "Convert to Sign Language" button.

[1436] 2. The device converts the entered lyrics into a JSON format request and sends it to the server using the HTTP protocol.

[1437] 3. The server analyzes the lyrics text and generates sign language gesture data using a generative AI model.

[1438] 4. The server generates a movement sequence for the avatar based on the sign language gesture data, and the emotion engine recognizes the user's emotions and reflects them in the movement.

[1439] 5. The server generates the final avatar dance video and uploads it to a social media platform.

[1440] 6. The user watches the video on a social media platform and checks the meaning of the sign language.

[1441] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1442] Step 1:

[1443] The user enters the lyrics text.

[1444] Users access the system through a web browser, paste the music lyrics into the input form, and then click the "Convert to Sign Language" button.

[1445] Input: Lyric text (e.g. "Haru yo, Koi" (Come Spring))

[1446] Output: Signals the start of a request to the system

[1447] Step 2:

[1448] The terminal transmits the lyrics text data to the server.

[1449] The device receives the lyrics data entered by the user, converts it into a JSON-formatted request, and then sends the request to the server using the HTTP protocol.

[1450] Input: Lyric text

[1451] Data processing: Convert lyrics data into JSON format

[1452] Output: HTTP request to the server

[1453] Step 3:

[1454] The server parses the lyrics text.

[1455] The server parses the received JSON request and extracts the lyrics text, then uses a natural language processing library (e.g., spaCy) to analyze the lyrics and extract important keywords and phrases.

[1456] Input: Lyrics data in JSON format

[1457] Data processing: Lyric text analysis, keyword extraction

[1458] Output: Analyzed text data (keywords and phrases)

[1459] Step 4:

[1460] The server generates sign language gesture data.

[1461] The server uses a generative AI model (e.g., a TensorFlow model) to convert the parsed text data into sign language gesture data.

[1462] Input: Parsed text data

[1463] Data Computation: Transformation with Generative AI Models

[1464] Output: Sign language gesture data

[1465] Step 5:

[1466] The server generates the avatar movement sequence.

[1467] The server generates avatar movement sequences based on sign language gesture data using 3D content creation tools such as Blender and Unity, and analyzes music files to combine movement sequences into a timeline in sync with the rhythm and tone of the music.

[1468] Input: Sign language gesture data, music files

[1469] Data processing: generating motion sequences and combining timelines

[1470] Output: Avatar movement sequence

[1471] Step 6:

[1472] The server uses an emotion engine to recognize the user's emotion.

[1473] The server uses facial recognition technology (e.g., OpenCV, dlib) and voice analysis libraries (e.g., Google Cloud Speech-to-Text API) to detect the user's emotional state in real time. The recognized emotional information is reflected in the avatar's behavior.

[1474] Input: User's face image, voice data

[1475] Data Computing: Applying Emotion Recognition Algorithms

[1476] Output: User's emotional state information

[1477] Step 7:

[1478] The server generates a dance video of the avatar.

[1479] The server generates a dance video of the avatar based on the final movement sequence. It renders the avatar's movements using Blender or Unity and exports them to a video file (e.g., MP4 format) using FFmpeg.

[1480] Input: Action sequence, emotional state information

[1481] Data processing: video rendering, file export

[1482] Output: Dance video file

[1483] Step 8:

[1484] The server distributes the generated video over the network.

[1485] The server uploads the generated video file to a social media platform (e.g. YouTube, Instagram), allowing users to access the video file.

[1486] Input: Dance video file

[1487] Data transmission: Video upload

[1488] Output: Public videos on social media platforms

[1489] Step 9:

[1490] A user watches a video on a social media platform.

[1491] Users access the social media platform and watch videos of the generated avatar dancing, with associated metadata such as the original lyrics and sign language meanings displayed alongside the video.

[1492] Input: Access to social media platforms

[1493] Output: Watchable dance video and associated metadata

[1494] (Application example 2)

[1495] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1496] Conventional sign language translation systems cannot reflect the user's emotions when converting music lyrics into sign language and expressing them with an avatar, making it difficult to add emotional expression to the sign language movements or avatar dance. Furthermore, even if the generated video content is visually rich, it lacks movements that correspond to the user's emotions, limiting the viewing experience.

[1497] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1498] In this invention, the server includes means for converting music lyrics into sign language, means for generating avatar movements based on the sign language, means for generating a video by synchronizing the generated avatar movements with the rhythm of the music, means for recognizing the user's emotions and adding emotional expressions to the avatar movements, and means for distributing the generated video via a network. This makes it possible to generate a sign language dance video that incorporates expressions according to the user's emotions, providing a richer viewing experience.

[1499] "Music" is an art form that expresses human emotion and experience in the form of sound, and includes rhythm, melody, and harmony.

[1500] "Lyrics" are a part of music that expresses the content and theme of a song in words.

[1501] "Sign language" is a visual, gestural language used for communication between people with hearing impairments, using hand and finger movements and facial expressions.

[1502] An "avatar" refers to a virtual person or character that symbolically represents a user or character in a digital space.

[1503] "Movement" refers to the physical movements and gestures performed by the avatar, including things like sign language and dancing.

[1504] "Video" is a media format that visually recreates movement by displaying a series of still images in succession.

[1505] "Rhythm" refers to the pattern of changes in the length and dynamics of notes in music, and is an element that gives a piece of music a sense of unity and movement.

[1506] "Emotions" refer to the psychological states that humans feel in response to specific situations or events, and include joy, sadness, anger, etc.

[1507] A "network" is a system of computers and communication devices connected together to share information and data.

[1508] "Distribution" refers to the act of providing or transmitting digital content to users over a network.

[1509] A "server" is a computer system that provides data and services over a network.

[1510] A "generative AI model" is a model generated by artificial intelligence and trained to predict or generate appropriate outputs for specific inputs.

[1511] A "prompt sentence" is an input sentence provided to a generative AI model to obtain a desired output.

[1512] This system converts music lyrics into sign language and generates and distributes videos of avatars dancing with the sign language. This system also incorporates a function to recognize the user's emotions and add emotional expressions to the avatar's movements.

[1513] System Configuration

[1514] The system consists of the following main components: a user device, a server, a generative AI model, an emotion recognition engine, and a video distribution platform.

[1515] Program processing description

[1516] 1. User Device

[1517] - Lyrics text input: The user device provides an input form for entering music lyrics. The user enters the lyrics text of their favorite song into this form and clicks the "Convert to sign language" button.

[1518] - Emotion recognition: The user's face is captured in real time through the camera on the user's device, and the emotion recognition engine analyzes the user's emotions.

[1519] 2. Data Transmission

[1520] - The device converts the entered lyrics data into JSON format and sends it to the server. At the same time, the user's emotional data is also sent.

[1521] 3. Server Processing

[1522] - The server analyzes the received lyrics text and converts them into sign language gesture data using a generative AI model (e.g., OpenAI's GPT-4). This conversion process uses prompts.

[1523] - Use an emotion recognition engine (e.g., Microsoft Azure Emotion Recognition API) to obtain the user's emotional information and reflect it in the avatar's behavior.

[1524] 4. Avatar Motion Generation

[1525] The server generates a movement sequence for the avatar based on the generated sign language gesture data, which specifically instructs the movements of each joint and body part of the avatar.

[1526] - Perform rhythmic analysis of the music (e.g., librosa library) and synchronize the avatar's movement sequences to the music.

[1527] 5. Video Generation

[1528] - The server generates a dance video by combining the avatar's movement sequence with music, using a video generation tool (e.g., OpenCV and ffmpeg library). The generated video is saved in MP4 format.

[1529] 6. Video streaming

[1530] - The server distributes the generated videos over the network, and uploads them to social media platforms (e.g. YouTube API, Facebook API) so that users can easily watch and share them within the app.

[1531] Examples of concrete examples and prompts

[1532] As a concrete example, when the lyrics "Haru yo, Koi" (Come Spring) are entered by a user, the generative AI model generates sign language gesture data using the following prompt sentence:

[1533] Input lyrics:

[1534] "Come Spring"

[1535] Prompt statement:

[1536] Convert the following lyrics into sign language and generate sign language gesture data for the avatar. Also, please consider that the user's emotion is "joy." Output the generated gesture data in JSON format.

[1537] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1538] Step 1:

[1539] The user inputs music lyric text into the terminal and clicks the "Convert to Sign Language" button.

[1540] Input: Music lyric text

[1541] Output: Lyric text input confirmation

[1542] First, the user enters the lyrics of the music into the application's dedicated input form and clicks the "Convert to Sign Language" button. At this time, the user's camera is activated and captures facial images in real time, which prepares the emotion recognition engine to analyze the user's emotions.

[1543] Step 2:

[1544] The device converts the lyrics text into JSON format and sends it to the server, along with the user's facial image data.

[1545] Input: Lyric text, user's facial image data

[1546] Output: Sending lyrics data in JSON format and facial video data to the server

[1547] The device converts the entered lyrics text into JSON format and sends it to a cloud server. A video of the user's face is also transferred at the same time, allowing subsequent processing to proceed smoothly.

[1548] Step 3:

[1549] The server analyzes the lyric text and generates sign language gesture data using a generative AI model.

[1550] Input: Lyrics data in JSON format

[1551] Output: Sign language gesture data

[1552] The server parses the received JSON-formatted lyrics data and converts the lyrics text into sign language gesture data using a generative AI model (e.g., OpenAI's GPT-4). During this process, an appropriate prompt is passed to the generative AI model.

[1553] Step 4:

[1554] The server uses an emotion recognition engine to analyze the user's emotions and obtain emotion data.

[1555] Input: Facial image data

[1556] Output: Emotion data

[1557] The server uses an emotion recognition engine (e.g., Microsoft Azure Emotion Recognition API) to analyze the facial video data and identify the user's emotions. The emotion data is saved so that it can be reflected in the avatar's behavior.

[1558] Step 5:

[1559] The server generates an avatar's movement sequence based on the sign language gesture data and emotion data.

[1560] Input: Sign language gesture data, emotion data

[1561] Output: Avatar movement sequence

[1562] The server combines the sign language gesture data and emotion data to generate a movement sequence for the avatar, which generates realistic and emotional avatar movements based on the user's emotions. Specific instructions for the movements of each joint and body part are created.

[1563] Step 6:

[1564] The server analyzes the rhythm of the music and synchronizes the avatar's movement sequence with the music.

[1565] Input: Action sequences, music files

[1566] Output: rhythm-synchronized movement sequence

[1567] The server analyzes the music files to identify the rhythm (e.g., using the librosa library), then combines the rhythmic timeline with a movement sequence to create engaging, music-synchronized avatar movements.

[1568] Step 7:

[1569] The server combines the avatar's movement sequence with music to generate a video.

[1570] Input: rhythm-synchronized movement sequences, music files

[1571] Output: Generated video file

[1572] The server finally generates a dance video by combining the avatar's movement sequence with music. This video generation uses libraries such as OpenCV and ffmpeg to generate a video file in MP4 format.

[1573] Step 8:

[1574] The server distributes the generated video to the social media platform via the network.

[1575] Input: Generated video file

[1576] Output: Upload to social media platforms

[1577] The server uploads the generated video to a social networking platform (e.g., YouTube API, Facebook API) for users to easily access, watch, and share, allowing users and viewers to enjoy the emotional avatar movements and music.

[1578] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1579] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1580] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1581] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1582] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1583] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1584] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1585] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1586] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1587] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1588] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1589] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1590] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1591] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1592] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1593] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1594] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1595] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1596] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1597] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1598] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1599] The following is further disclosed regarding the above embodiment.

[1600] (Claim 1)

[1601] A means of converting music lyrics into sign language,

[1602] means for generating avatar movements based on sign language;

[1603] A means for generating a video by synchronizing the movements of the generated avatar with the rhythm of music;

[1604] A means for distributing the generated video over a network;

[1605] A system including:

[1606] (Claim 2)

[1607] a means for inputting text data of lyrics;

[1608] The means for converting into sign language uses an AI model that analyzes text data of lyrics and converts it into sign language gesture data.

[1609] 10. The system of claim 1.

[1610] (Claim 3)

[1611] The means for generating the avatar's movement is characterized by using a script that analyzes the sign language gesture data and generates a joint movement sequence of the avatar.

[1612] 10. The system of claim 1.

[1613] "Example 1"

[1614] (Claim 1)

[1615] a means for inputting text data of lyrics;

[1616] A means for converting input lyrics text data into JSON format and sending it to the server;

[1617] A means for analyzing lyric text data using an AI model and converting it into sign language gesture data;

[1618] means for generating a movement sequence of an avatar based on sign language gesture data;

[1619] A means for analyzing music files and synchronizing avatar movement sequences with the rhythm of the music;

[1620] means for generating a dance video based on a movement sequence of an avatar;

[1621] The generated video can be saved in a common video format such as MP4 and distributed via a network.

[1622] A means for viewing videos distributed via a network;

[1623] A system including:

[1624] (Claim 2)

[1625] It features an AI model that analyzes lyrics text data and converts it into sign language gesture data.

[1626] 10. The system of claim 1.

[1627] (Claim 3)

[1628] It analyzes sign language gesture data and uses scripts that specifically instruct the movements of each joint and body part of the avatar.

[1629] 10. The system of claim 1.

[1630] "Application Example 1"

[1631] (Claim 1)

[1632] A means of converting music lyrics into sign language,

[1633] means for generating avatar movements based on sign language;

[1634] A means for generating a video by synchronizing the movements of the generated avatar with the rhythm of music;

[1635] A means for distributing the generated video over a network;

[1636] A means to convert the text of the recipe into sign language and generate a video using sign language by an avatar;

[1637] A means for distributing the generated sign language video to people with visual impairments;

[1638] A system including:

[1639] (Claim 2)

[1640] a means for inputting text data of lyrics;

[1641] The means for converting into sign language uses an AI model that analyzes text data of lyrics and converts it into sign language gesture data.

[1642] 10. The system of claim 1.

[1643] (Claim 3)

[1644] The means for generating the avatar's movement is characterized by using a script that analyzes the sign language gesture data and generates a joint movement sequence of the avatar.

[1645] 10. The system of claim 1.

[1646] "Example 2: Combining Emotion Engines"

[1647] (Claim 1)

[1648] A means of converting music lyrics into sign language,

[1649] means for generating avatar movements based on sign language;

[1650] A means for generating a video by synchronizing the movements of the generated avatar with the rhythm of music;

[1651] A means for distributing the generated video over a network;

[1652] means for recognizing a user's emotions and adding emotional expressions to the avatar's movements;

[1653] A system including:

[1654] (Claim 2)

[1655] a means for inputting text data of lyrics;

[1656] The means for converting into sign language uses a generative AI model that analyzes text data of lyrics and converts it into sign language gesture data.

[1657] 10. The system of claim 1.

[1658] (Claim 3)

[1659] The means for generating the avatar's movement is characterized by using a script that analyzes the sign language gesture data and generates a joint movement sequence of the avatar.

[1660] 10. The system of claim 1.

[1661] "Application example 2 when combining emotion engines"

[1662] (Claim 1)

[1663] A means of converting music lyrics into sign language,

[1664] means for generating avatar movements based on sign language;

[1665] A means for generating a video by synchronizing the movements of the generated avatar with the rhythm of music;

[1666] A means for distributing the generated video over a network;

[1667] and means for recognizing a user's emotions and adding emotional expressions to the avatar's movements.

[1668] (Claim 2)

[1669] a means for inputting text data of lyrics;

[1670] The means for converting into sign language uses a generative AI model that analyzes text data of lyrics and converts it into sign language gesture data.

[1671] 10. The system of claim 1.

[1672] (Claim 3)

[1673] The means for generating the avatar's movement is characterized by using a script that analyzes the sign language gesture data and generates a joint movement sequence of the avatar.

[1674] 10. The system of claim 1. [Explanation of symbols]

[1675] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of converting music lyrics into sign language, means for generating avatar movements based on sign language; A means for generating a video by synchronizing the movements of the generated avatar with the rhythm of music; A means for distributing the generated video over a network; A system including:

2. a means for inputting text data of lyrics; The means for converting into sign language uses an AI model that analyzes text data of lyrics and converts it into sign language gesture data. The system of claim 1 .

3. The means for generating the avatar's movement is characterized by using a script that analyzes the sign language gesture data and generates a joint movement sequence of the avatar. The system of claim 1 .

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A