system

A system using generative AI to provide sign language videos and volunteer information addresses the interpreter shortage, facilitating effective communication and volunteer participation.

JP2026063734APending Publication Date: 2026-04-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-10-01
Publication Date
2026-04-13

AI Technical Summary

Technical Problem

The shortage of sign language interpreters at events and international conferences, particularly in multilingual settings, hinders effective communication and volunteer participation, with a lack of resources for learning sign language and promoting volunteer activities.

Method used

A system that allows users to input sign language words via a terminal, generates corresponding videos using generative artificial intelligence, and provides volunteer activity information, facilitating learning and participation.

Benefits of technology

Efficiently supports sign language learning and promotes volunteer activities, addressing the interpreter shortage and enhancing communication in multilingual environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026063734000001_ABST
    Figure 2026063734000001_ABST
Patent Text Reader

Abstract

This system aims to address the shortage of sign language interpreters while simultaneously achieving multilingual support and promoting volunteer activities. [Solution] A system comprising: means for the user to input into a search box via a terminal; means for the terminal to generate a search request and send it to a server; means for the server to receive the search request and extract the entered word; means for the server to search for the corresponding sign language word in a database; means for the server to call a generative artificial intelligence and generate a sign language video; means for the server to return the search results to the user's terminal; means for the terminal to process the response received from the server and trim and play the sign language video; means for the user to watch the sign language video on the terminal and learn sign language; and means for the terminal to display volunteer activity information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] An object of the present invention is to address the problem of a shortage of sign language interpreters at events for the hearing-impaired and international conferences. In particular, in situations where multilingual support is required, the number of international sign language interpreters is insufficient, and there is a need for means that allow users who want to learn sign language to learn efficiently. In addition, it is also necessary to provide a support tool for people who do not know sign language to communicate smoothly with foreigners. Furthermore, it is also a problem to encourage many people to participate in volunteer activities. p

Means for Solving the Problems

[0005] This invention provides a system in which a user enters a sign language word into a search box via a terminal, generates a search request, and sends it to a server. The server receives the request, extracts the entered word, and searches for the corresponding sign language word in its database. If the corresponding sign language word does not exist in the database, it invokes a generative artificial intelligence to generate a sign language video. The generated sign language video is sent back from the server to the user's terminal and streamed on the terminal. The user can learn sign language through this and, after learning, can receive volunteer activity information from the terminal. This can compensate for the shortage of sign language interpreters and simultaneously achieve multilingual support and the promotion of volunteer activities.

[0006] A "user" is an individual or group that uses the system, with the purpose of learning sign language or participating in volunteer activities.

[0007] A "terminal" refers to a device operated by a user, such as a smartphone or PC.

[0008] A "search box" is an interface element for users to input sign language words; it is an input field displayed on the device.

[0009] A "search request" is data containing sign language words entered by the user, and it is an HTTP request sent from the terminal to the server.

[0010] A "server" is a computer system that receives search requests and performs tasks such as searching for sign language words or calling up generative artificial intelligence.

[0011] A "database" is a collection of sign language words and their corresponding image and video data that are managed and stored, and it is the information source that a server uses to perform queries and searches.

[0012] "Generative artificial intelligence" is an AI technology used to generate videos and images of sign language words requested by users, and is used to supplement sign language data that does not exist in the database.

[0013] A "sign language video" is video data that records hand movements and facial expressions that represent specific sign language words or phrases.

[0014] "Response data" refers to the search result data that the server sends back to the terminal, including the URL of the sign language video and other related information.

[0015] "Streaming playback" is a method of playing sign language videos on a user's device by receiving them sequentially from the internet.

[0016] "Volunteer activity guidance" refers to the display of information about volunteer activities provided by organizations like SoftBank on the device of users who have completed sign language learning. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8]It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the language used in the following description will be described.

[0020] In the following embodiments, a processor with a reference number (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Further, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of the arithmetic unit include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] This invention provides a system that makes it easier for users to learn sign language and also offers tools to promote participation in volunteer activities. The system searches for sign language words entered through a search box and supports learning through images and videos. The system's program processing flow is described below in natural language, with detailed examples.

[0039] System processing flow

[0040] 1. User search for sign language words:

[0041] The user opens the application on their device and enters the sign language word they want to learn into the search box. For example, they might type "hello" and click the search button.

[0042] 2. Sending request data via the terminal:

[0043] The terminal converts the search request containing the words entered by the user into JSON format and sends it to the server using the HTTP POST method. For example, the request body {"query": "Hello"} is generated and sent to the server.

[0044] 3. Receiving and parsing requests by the server:

[0045] The server receives the HTTP request and extracts the search word from the request body. Based on the extracted word "hello," a query is generated to search for the corresponding sign language word in the database.

[0046] 4. Database search by server:

[0047] The server searches the database for the relevant sign language word. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If it does not exist, a generative artificial intelligence is called to generate a sign language video.

[0048] 5. Server-based invocation of generative artificial intelligence:

[0049] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video corresponding to "hello." This generated video is stored in the database as a cache for future searches.

[0050] 6. Server sends response data:

[0051] The server constructs the search result response data and sends it back to the terminal. This response data includes the URL of a sign language video for "hello."

[0052] 7. Processing and display of response data by the terminal:

[0053] The terminal receives the response sent back from the server and extracts the URL of the sign language video. Based on this, it initializes the video player and streams the sign language video to the user.

[0054] 8. User learning of sign language:

[0055] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. For example, a user might practice the sign for "hello" while watching a video.

[0056] 9. Volunteer activity guidance via terminal:

[0057] After watching the video, the device will automatically display information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" will appear on the device.

[0058] 10. User registration for volunteer activities:

[0059] Users follow the instructions and register to participate in volunteer activities that interest them. For example, a user can participate in an activity by filling in the required information on a form and clicking the registration button.

[0060] Thus, the system of the present invention efficiently supports sign language learning and motivates users to participate in volunteer activities. As a result, it can help compensate for the shortage of sign language interpreters and contribute to the promotion of social contribution activities.

[0061] The following describes the processing flow.

[0062] Step 1:

[0063] The user enters the sign language word they want to learn into the search box via their device. After entering the word, they click the search button. For example, they might enter "hello" and press the search button.

[0064] Step 2:

[0065] The terminal receives user input data and generates a search request in JSON format. The generated request is sent to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[0066] Step 3:

[0067] The server receives an HTTP request. It extracts the search term "hello" from the received request body and generates the query necessary for the database search.

[0068] Step 4:

[0069] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[0070] Step 5:

[0071] The server checks the search results from the database. If a matching sign language video exists in the database, it retrieves the URL of that video. If it does not exist, it proceeds to the next step.

[0072] Step 6:

[0073] The server invokes a generative artificial intelligence to generate a sign language video corresponding to the sign language word "hello" requested by the user. This generated video is stored as a cache in the database as feedback.

[0074] Step 7:

[0075] The server constructs response data that includes the URL of the sign language video. For example, it generates data in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[0076] Step 8:

[0077] The server sends the generated response data back to the terminal. The terminal receives the HTTP response.

[0078] Step 9:

[0079] The terminal processes the response received from the server and extracts the URL of the sign language video. Using the extracted URL, the video player is initialized and the sign language video is streamed.

[0080] Step 10:

[0081] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements. Specifically, they practice the sign language for "hello" while watching the video.

[0082] Step 11:

[0083] After watching a sign language video, the device automatically displays information about volunteer activities. For example, a pop-up message might appear asking, "Would you like to participate in volunteer activities?"

[0084] Step 12:

[0085] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0086] (Example 1)

[0087] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0088] Traditional sign language learning systems often lacked relevant video resources when users wanted to learn specific sign language words, which acted as a barrier to learning. Furthermore, they lacked means to motivate participation in volunteer activities in addition to promoting sign language learning. Therefore, there was a need for a system that could simultaneously provide efficient sign language learning support and promote volunteer activities.

[0089] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0090] In this invention, the server includes means for calling a generative artificial intelligence to generate sign language videos, means for storing the sign language videos generated using the generative artificial intelligence as a cache in a database, and means for the server to store search results as a cache in the database. This allows for the real-time generation of sign language videos when a user searches for a specific sign language word, and further enables faster query responses for the same search in the future. In addition, by simultaneously providing information on volunteer activities during the sign language learning process, participation in volunteer activities can also be promoted.

[0091] "User" refers to an individual who uses the sign language learning system.

[0092] A "device" refers to an electronic device used by a user, such as a smartphone, tablet, or personal computer.

[0093] A "search box" refers to the interface used by users to input sign language words they want to learn.

[0094] A "search request" refers to data sent from the device to the server based on the words the user enters into the search box.

[0095] A "server" refers to the central processing unit of a sign language learning system, which is responsible for accessing the database and calling generative artificial intelligence.

[0096] A "database" refers to an information management system that stores sign language words and their associated video data.

[0097] "Generative artificial intelligence" refers to machine learning models and algorithms used to generate sign language videos corresponding to specific sign language words.

[0098] A "sign language video" refers to a video or image that visually demonstrates specific sign language words.

[0099] "Cache" refers to a temporary storage area for data such as generated sign language videos, in order to improve the response speed to future search queries.

[0100] "Streaming playback" refers to the real-time playback of sign language videos on the user's device.

[0101] "Volunteer Activity Guide" refers to information provided to users through the sign language learning system to encourage participation in social contribution activities.

[0102] This invention is a system designed to make it easier for users to learn sign language and to further promote participation in volunteer activities. The system has a mechanism that allows users to input specific sign language words through a search box, and then provides a corresponding sign language video. The specific implementation method of the system is described below.

[0103] Hardware and software configuration

[0104] User terminal

[0105] Users access the system using devices such as smartphones, tablets, or personal computers. These devices require an internet connection, and the system is accessed through a browser application or a dedicated application.

[0106] server

[0107] The server is the central component for receiving and processing user requests. The server implements the following software and functions:

[0108] Web server: Software used to receive HTTP requests and return responses.

[0109] JSON parser: A library for analyzing the format of received data.

[0110] Database Management System: A system for storing and searching sign language words and their corresponding sign language videos.

[0111] Generative artificial intelligence models: Machine learning models and algorithms for generating sign language videos in real time.

[0112] Flow of operations

[0113] 1. User search for sign language words

[0114] The user opens the application on their device and enters the sign language word they want to learn into the search box. For example, they might enter "hello" and click the search button. This sends a search request to the system.

[0115] 2. Sending request data by the terminal

[0116] The terminal converts the search request containing the words entered by the user into JSON format and sends it to the server using the HTTP POST method. Specifically, the generated request data is {"query": "こんにちは"}.

[0117] 3. Receiving and parsing requests by the server

[0118] The server receives the HTTP request and extracts the search word from the request body. Based on the extracted word "hello," a query is generated to search for the corresponding sign language word in the database.

[0119] 4. Server-based database searches and calls to generative artificial intelligence.

[0120] The server searches the database for the relevant sign language word. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If no such data exists in the database, the server calls a generative artificial intelligence to generate a sign language video corresponding to "hello." An example of a prompt to the generative AI model is, "Please generate a sign language video for 'hello'."

[0121] 5. Server sends response data

[0122] The server constructs the search result response data and sends it back to the terminal. This response data includes the URL of a sign language video of "hello." For example, it is sent in the form of {"videoURL": "http: / / example.com / videos / kon-nichiwa.mp4"}.

[0123] 6. Processing and display of response data by the terminal.

[0124] The device receives the response sent back from the server and extracts the URL of the sign language video. Using this URL, the device initializes its video player and streams the sign language video to the user.

[0125] 7. User learning of sign language

[0126] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. For example, a user might practice the sign for "hello."

[0127] 8. Display of volunteer activity information on terminals

[0128] After watching the video, the device automatically displays information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" might appear.

[0129] 9. User registration for volunteer activities

[0130] Users follow the instructions and register to participate in volunteer activities that interest them. For example, they can participate in an activity by filling in the required information on a form and clicking the registration button.

[0131] As described above, this system not only efficiently supports sign language learning but also promotes users' participation in volunteer activities. This helps to alleviate the shortage of sign language interpreters and contributes to the promotion of social contribution activities.

[0132] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0133] Step 1:

[0134] The user launches the application on their device and enters the sign language word they want to learn into the search box. For example, the user enters "hello" and clicks the search button. The device retrieves the contents of the search box based on the user's input and generates JSON data {"query": "hello"}.

[0135] Step 2:

[0136] The terminal sends the generated JSON-formatted search request to the server using the HTTP POST method. Specifically, it uses an HTTP client library to send the request data to the specified endpoint on the server.

[0137] Input: Sign language word "hello" entered in the search box.

[0138] Data processing: Convert search requests to JSON

[0139] Output: Search request in JSON format {"query": "Hello"}

[0140] Step 3:

[0141] The server extracts the request body from the received HTTP request, parses the JSON data to extract the search term "hello". Based on this, the server generates an SQL query SELECT FROM SignLanguageVideos WHERE keyword='hello' to search for the corresponding sign language word in the database.

[0142] Input: HTTP request body

[0143] Data processing: JSON parsing and extraction of search terms.

[0144] Output: SQL query SELECT FROM SignLanguageVideos WHERE keyword='こんにちは'

[0145] Step 4:

[0146] The server executes the generated SQL query against the database to search for the corresponding sign language video. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If it does not exist, a generative artificial intelligence is called to generate a sign language video.

[0147] Input: SQL query

[0148] Data Calculation: Executing Database Queries

[0149] Output: URL of the sign language video or status of not found.

[0150] Step 5:

[0151] If the server does not find a corresponding sign language word in the database, it uses a generative artificial intelligence model to generate a sign language video corresponding to "hello." A prompt message, "Please generate a sign language video for 'hello'," is sent to the generative AI model, and the generated video file is received.

[0152] Input: Search term "hello" - Status: Not found

[0153] Data processing: Generate prompt sentences, send them to an AI model, and generate a video.

[0154] Output: Generated sign language video file

[0155] Step 6:

[0156] The generated sign language videos are stored as a cache in the database to enable faster responses to future queries. This eliminates the need to generate videos again when a search is performed.

[0157] Input: Generated sign language video file

[0158] Data processing: Saving video files

[0159] Output: Sign language video files stored in the database

[0160] Step 7:

[0161] The server constructs response data containing the URL of the generated sign language video or a sign language video retrieved from the database, and sends it back to the terminal. For example, it generates a JSON response such as {"videoURL": "http: / / example.com / videos / kon-nichiwa.mp4"}.

[0162] Input: URL of a sign language video or a generated sign language video

[0163] Data processing: Building response data

[0164] Output: Response data in JSON format

[0165] Step 8:

[0166] The device receives the response sent back from the server and extracts the URL of the sign language video from it. Using the extracted URL, the device initializes the video player and streams the sign language video.

[0167] Input: Response data in JSON format

[0168] Data processing: Parsing responses and extracting URLs.

[0169] Output: Initialization of the video player and streaming playback

[0170] Step 9:

[0171] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. Specifically, they practice the sign for "hello" while watching the videos.

[0172] Input: Watching sign language videos

[0173] Data calculation: Imitation of hand movements

[0174] Output: Learning sign language

[0175] Step 10:

[0176] After watching a video, the device automatically displays information about volunteer activities to the user. For example, it might display a pop-up asking, "Would you like to participate in volunteer activities?"

[0177] Input: Video viewing end event

[0178] Data processing: Displaying volunteer guidance messages

[0179] Output: Display of a popup message

[0180] Step 11:

[0181] Users follow the instructions and register to participate in volunteer activities that interest them. Specifically, they participate in an activity by filling in the required information on a form and clicking the registration button.

[0182] Input: User-generated registration information

[0183] Data processing: Form data submission

[0184] Output: Registration for volunteer activities

[0185] Through the steps outlined above, this system efficiently supports sign language learning and promotes users' participation in volunteer activities.

[0186] (Application Example 1)

[0187] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0188] In modern factory and other work environments, communication between hearing-impaired workers and machines or robots remains a challenge. Furthermore, in environments where understanding sign language is insufficient, smooth communication using sign language is difficult, potentially reducing the work efficiency of hearing-impaired workers. Additionally, a lack of motivation among the general public to learn sign language and participate in volunteer activities is also a problem.

[0189] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0190] In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in a database, means for the server to invoke a generative artificial intelligence to generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to display volunteer activity information, means for capturing the user's sign language movements using a camera and analyzing them, means for sending the analyzed sign language movements to the server and recognizing the instruction content, and means for the robot to perform an action based on the recognized instruction content. This enables smooth communication between workers with hearing impairments and robots in the factory, and makes it possible to promote sign language learning and motivate participation in volunteer activities.

[0191] A "search box" is an input field used by users to enter specific words or phrases via their device.

[0192] A "search request" is a request generated by the device based on data entered by the user and sent to the server.

[0193] A "server" is a computer system that receives and processes search requests.

[0194] A "database" is an information system for systematically storing and managing data such as sign language words and their corresponding videos, in a searchable format.

[0195] "Generative artificial intelligence" refers to artificial intelligence algorithms that can generate new sign language videos based on existing data.

[0196] A "sign language video" is a video recording of actions that express specific sign language words or phrases.

[0197] "Streaming playback" is a method of playing video data in real time, allowing users to watch before the data is fully downloaded.

[0198] A "volunteer activity guide" is a notification or message that provides users with information about volunteer activities and encourages them to participate.

[0199] A "camera" is an imaging device used to capture the user's sign language movements.

[0200] "Analysis" is the process of deciphering sign language movements captured by a camera and determining whether they represent specific sign language words or phrases.

[0201] A "robot" is a mechanical device that autonomously performs specific tasks in factories and other settings, operating based on user instructions.

[0202] A "user" is a person who uses this system to learn sign language or to give instructions to the robot through sign language.

[0203] This invention provides a system that makes it easier for users to learn sign language and enables smooth communication with robots using sign language within a factory. This system is primarily implemented using a user terminal, server, database, camera, and robot.

[0204] User terminal functions

[0205] User devices include smartphones, tablets, and personal computers. These devices are equipped with a search box where users enter sign language words or phrases they wish to learn. A search request is generated and sent to the server. The device has the capability to convert the request into JSON format.

[0206] Server Functions

[0207] The server receives a search request from the user and extracts the words from the request. Next, it searches the database for videos of the corresponding sign language words. If no video exists in the database, it calls upon a generative artificial intelligence to generate a sign language video. The generated video is also stored in the database as a cache for future searches. The server returns the search results to the user's terminal and provides the URL of the sign language video.

[0208] Camera functions

[0209] The camera is used to capture the user's sign language movements. The captured video is analyzed by an internal or external sign language recognition algorithm, and the content of the sign language is extracted. The extracted content is then sent back to the server to be used as instructions for the robot's actions.

[0210] Robot functions

[0211] The robots in the factory perform actions based on sign language instructions received from a server. For example, if a user signs "start work," the robot recognizes the sign and begins the specified task.

[0212] Specific Examples and Generative AI Models

[0213] As a concrete example of this system's use, consider a scenario where a user gives the sign language instruction to "stop work." A camera captures the sign language movement, analyzes the data, and sends it to a server. The server analyzes the data and sends the corresponding instruction to the robot. As a result, the robot immediately stops working. Additionally, when a user searches for sign language videos to learn, the generative artificial intelligence generates new sign language videos and saves them to the database.

[0214] Examples of prompt sentences to use for sign language recognition in a generative AI model are as follows:

[0215] "Design and implement a method to recognize specific sign language actions (e.g., 'start work') from short video clips."

[0216] This system will allow hearing-impaired employees to learn sign language and give natural sign language instructions to robots within the factory, which is expected to improve the working environment and increase operational efficiency.

[0217] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0218] Step 1:

[0219] The user enters the sign language word they want to learn into the search box via their device. The entered word is sent to the device as request data. Specifically, the user enters "hello" and presses the search button. In this case, the input is "hello," and it is recognized by the device.

[0220] Step 2:

[0221] The device generates a search request, converts it to JSON format, and sends it to the server. The generated JSON data is in the format {"query": "こんにちは"} and is sent to the server using the HTTP POST method. The input in this case is the word "こんにちは" entered by the user, and the output is data in JSON format.

[0222] Step 3:

[0223] The server receives a search request and extracts a word from the request body. Specifically, it extracts the value "こんにちは" from the "query" field of the received JSON data. The input in this case is JSON data {"query": "こんにちは"}, and the output is the extracted word "こんにちは".

[0224] Step 4:

[0225] The server searches the database for the corresponding sign language word. Specifically, it uses a database query to search for a sign language video corresponding to "hello." The input in this case is the extracted word "hello," and the output is the URL of the corresponding sign language video.

[0226] Step 5:

[0227] If the server does not find a corresponding sign language video in the database, it invokes a generative artificial intelligence to generate one. The generated sign language video is then saved as a cache in the database. The input in this case is the extracted word "hello," and the output is the generated sign language video.

[0228] Step 6:

[0229] The server returns the search results to the user's device. The search results include URLs for sign language videos. Specifically, the URLs are constructed as a JSON response and sent to the user's device. The input in this process is the search results or the URLs of the generated sign language videos, and the output is the JSON response sent to the user's device.

[0230] Step 7:

[0231] The terminal processes the response received from the server, extracts the URL of the sign language video, and starts streaming playback. Specifically, it initializes the video player and plays the sign language video. The input in this process is a JSON response, and the output is the streaming playback of the sign language video.

[0232] Step 8:

[0233] Users learn sign language by watching sign language videos on their devices. Specifically, they imitate hand movements while watching the videos being played. In this process, the input is the sign language video, and the output is the learning of sign language.

[0234] Step 9:

[0235] The device uses its camera to capture the user's sign language movements and analyzes them. Specifically, it analyzes the video data captured by the camera using an internal algorithm to extract the content of the sign language. The input in this process is the captured sign language video, and the output is the extracted sign language content.

[0236] Step 10:

[0237] The terminal analyzes the sign language gestures and sends them to the server, which then recognizes the instructions. The analysis results are converted to JSON format and sent using the HTTP POST method. The input in this process is the extracted sign language content, and the output is the transmission of JSON data to the server.

[0238] Step 11:

[0239] The robot executes actions based on the instructions recognized by the server. Specifically, the robot receives recognition data from the server and performs the corresponding action, such as "start work." In this process, the input is the recognized instruction, and the output is the robot's action.

[0240] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0241] This invention relates to a system for promoting sign language learning and participation in volunteer activities, with the aim of supporting people with hearing impairments. Furthermore, this invention aims to provide users with an optimal learning experience by combining it with an emotion engine that recognizes the user's emotions. The processing flow of the system's program is described below in natural language and illustrated in detail with specific examples.

[0242] System processing flow

[0243] 1. User search for sign language words:

[0244] The user enters the sign language word they want to learn into the search box via their device and clicks the search button. For example, they might enter "hello".

[0245] 2. Sending request data via the terminal:

[0246] The terminal receives user input data, generates a search request in JSON format, and sends it to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[0247] 3. Receiving and parsing requests by the server:

[0248] The server receives the HTTP request, extracts the search term "hello" from the request body, and generates the query necessary for the database search.

[0249] 4. Database search by server:

[0250] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[0251] 5. Server-based invocation of generative artificial intelligence:

[0252] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video. This generated video is stored in the database as a cache for future searches.

[0253] 6. Server sends response data:

[0254] The server constructs response data containing the URL of the sign language video and sends it back to the terminal. For example, it sends data like {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"} to the terminal.

[0255] 7. Processing and display of response data by the terminal:

[0256] The terminal processes the response received from the server, extracts the URL of the sign language video, initializes the video player, and streams the sign language video.

[0257] 8. User learning of sign language:

[0258] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements. For example, they can practice the sign for "hello" while watching a video.

[0259] 9. User emotion recognition by emotion engine:

[0260] The device captures the user's facial expressions through its camera, and an emotion engine analyzes those expressions to recognize the user's emotions. For example, if the user is smiling during the learning process, it will be recognized as a positive emotion.

[0261] 10. Adjusting the learning content:

[0262] Based on the user's emotions recognized by the emotion engine, the device adjusts the learning content. For example, if it detects a confused expression on the user's face, it will provide supplementary learning materials or explanations.

[0263] 11. Recommendation of appropriate sign language learning content:

[0264] The device recommends appropriate sign language learning content based on the user's emotional state. For example, if the user is in a positive emotional state, it will suggest more difficult sign language.

[0265] 12. Display of volunteer activity information on terminals:

[0266] After learning sign language, the device automatically displays information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" might appear.

[0267] 13. User registration for volunteer activities:

[0268] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0269] As described above, the system of the present invention efficiently supports sign language learning and provides a more effective learning experience by adjusting the learning content according to the user's emotions. Furthermore, it can enhance social contribution by promoting participation in volunteer activities.

[0270] The following describes the processing flow.

[0271] Step 1:

[0272] The user enters the sign language word they want to learn into the search box on their device and clicks the search button. For example, they might enter "hello".

[0273] Step 2:

[0274] The terminal receives user input data and generates a search request in JSON format. The generated JSON request is sent to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[0275] Step 3:

[0276] The server receives an HTTP request. After receiving it, the search word "こんにちは" is extracted from the request body, and a query necessary for database search is generated. Specifically, the request is analyzed to identify the word entered by the user.

[0277] Step 4:

[0278] Using the query generated by the server, the corresponding sign language word "こんにちは" is searched in the database. For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[0279] Step 5:

[0280] The server checks the search results from the database. If the corresponding sign language video exists in the database, the URL of that video is obtained. If not, proceed to the next step.

[0281] Step 6:

[0282] The server calls the generative artificial intelligence to generate a sign language video corresponding to the sign language word "こんにちは" requested by the user. This generated video is saved as a cache in the database for future searches.

[0283] Step 7:

[0284] The server constructs response data including the URL of the sign language video. For example, data in the form of {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"} is generated.

[0285] Step 8:

[0286] The server returns the response data constructed to the user's terminal. The terminal receives the HTTP response.

[0287] Step 9:

[0288] The terminal analyzes the response received from the server and extracts the URL of the sign language video. Then, it initializes the video player and prepares it for streaming playback of the sign language video.

[0289] Step 10:

[0290] Users watch sign language videos on their devices and learn sign language by imitating the hand movements. Specifically, they practice the sign language for "hello" while watching the videos.

[0291] Step 11:

[0292] The device uses its camera to capture the user's facial expressions, and an emotion engine analyzes those expressions to recognize the user's emotions. For example, if the user is confused, the system analyzes and understands that expression.

[0293] Step 12:

[0294] Based on the user's emotions recognized by the emotion engine, the device adjusts the learning content. Specifically, if the user is confused, it displays supplementary materials or additional explanations.

[0295] Step 13:

[0296] The device recommends appropriate sign language learning content based on the user's emotional state. For example, if the user is in a positive state, it will suggest more difficult sign language.

[0297] Step 14:

[0298] After watching a sign language video, the device automatically displays information about social activities and volunteer opportunities to the user. For example, a pop-up message might appear asking, "Would you like to participate in volunteer activities?"

[0299] Step 15:

[0300] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0301] (Example 2)

[0302] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0303] Traditional sign language learning systems had problems such as the inability for users to easily obtain videos corresponding to the sign language words they searched for, which reduced the efficiency of sign language learning. Furthermore, the lack of customization of learning content based on the user's emotional state meant that learning effectiveness was not maximized, potentially leading to decreased user motivation. In addition, there was a lack of means to encourage participation in volunteer activities, resulting in fewer opportunities to raise awareness of social contribution.

[0304] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in the database, means for the server to call a generative artificial intelligence and generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to capture the user's facial expression via a camera and for the emotion engine to recognize the user's emotion, means for the terminal to adjust the learning content based on the recognized emotion, and means for the terminal to display volunteer activity information. As a result, the user can efficiently proceed with learning sign language, adjust the learning content based on their emotional state, and further increase their motivation to participate in volunteer activities.

[0305] A "terminal" is a device that a user directly operates to input and display information.

[0306] A "server" is a device that processes data transmitted from a terminal and provides necessary information.

[0307] A "search box" is a graphical user interface element used by a user to input specific information.

[0308] A "search request" is a request when querying a server for necessary data based on the information input by a user.

[0309] "JSON format" is short for JavaScript (registered trademark) Object Notation and is a lightweight text data exchange format for structuring and representing data.

[0310] "HTTP POST method" is an HTTP request method used by a client to send resources to a server.

[0311] An "emotion engine" is software or a device for recognizing and analyzing a user's emotions.

[0312] A "sign language video" is a video recording of sign language movements and expressions.

[0313] "Generative artificial intelligence" is an artificial intelligence technology with the ability to generate new content based on a given prompt text.

[0314] A "prompt text" is text data input to a generative artificial intelligence to generate a specific response.

[0315] "Streaming playback" is a technology for playing data while receiving it in real time.

[0316] "Volunteer Activity Guide" refers to guide information displayed to users to obtain information about volunteer activities and participate in them.

[0317] Modes for carrying out the invention

[0318] This invention relates to a sign language learning system for supporting people with hearing impairments and a system for promoting participation in volunteer activities. The system aims to provide an optimal learning experience by recognizing the user's emotions. The following describes in detail the embodiments for specifically carrying out this invention.

[0319] The user enters the sign language word they want to learn into the search box via their device and clicks the search button. Upon receiving this input, the device generates a search request in JSON format and sends it to the server using the HTTP POST method. For example, if the user wants to learn the sign language word for "hello," they would search for "hello" in the input, and the device would generate JSON data {"query": "hello"} and send it to the server.

[0320] The server receives this HTTP request and extracts the search word from the request body. Based on the extracted word, the server generates an SQL query to search for the corresponding sign language word in the database and executes the database search. For example, it executes the SQL query SELECT FROM sign_language WHERE word='こんにちは';

[0321] If the database search does not find a matching sign language word, the server calls a generative artificial intelligence (generative AI model) and generates a sign language video based on the prompt text. This generated sign language video is stored in the database as a cache for future searches. For example, the generative AI model will generate a sign language video if the prompt text "Please generate a sign language video for 'hello'" is entered into it.

[0322] A response data containing the URL of the generated or retrieved sign language video is constructed, and the server returns it to the terminal in JSON format. For example, it might be returned in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[0323] The device processes the received response data, extracts the URL of the sign language video, and initializes the video player. It then streams the sign language video, allowing the user to watch and learn from it on their device.

[0324] Furthermore, the device captures the user's facial expressions via the camera, and an emotion engine analyzes this facial data to recognize the user's emotions. For example, if the user is learning sign language with a smile, a positive emotion will be recognized. Based on this emotion data, the device adjusts the learning content. For instance, if the user shows a confused expression, the device will provide additional learning materials or clearer explanations.

[0325] After completing sign language learning, the device automatically displays information about volunteer activities and encourages the user to participate. For example, it might display a pop-up asking, "Would you like to participate in volunteer activities?" The user can then register to participate by following this prompt.

[0326] This system allows users to learn sign language efficiently, adjust learning content according to their emotions, and foster a sense of social responsibility by encouraging participation in volunteer activities.

[0327] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0328] Step 1:

[0329] The user enters the sign language word they want to learn into the search box via their device and clicks the search button.

[0330] Input: Sign language words entered by the user into the search box on the device (e.g., "hello")

[0331] Output: Clicking the search button sends the entered sign language word as a search request to the terminal.

[0332] Step 2:

[0333] The terminal receives user input data, generates a search request in JSON format, and sends it to the server using the HTTP POST method.

[0334] Input: Sign language word (e.g., "hello")

[0335] Data processing: Generate JSON data containing sign language words (e.g., {"query": "Hello"}).

[0336] Output: JSON data is sent to the server using the HTTP POST method.

[0337] Step 3:

[0338] The server receives an HTTP request and extracts the search term from the request body.

[0339] Input: JSON data sent to the server

[0340] Data processing: Extraction of search terms (e.g., "hello")

[0341] Output: Extracts search terms and prepares a query for database search.

[0342] Step 4:

[0343] The server searches the database for the corresponding sign language word.

[0344] Input: Search term (e.g., "hello")

[0345] Data calculation: Executing database queries (e.g., SELECT FROM sign_language WHERE word='こんにちは';)

[0346] Output: Search results (If a relevant sign language video is found, that will be the result; otherwise, proceed to the next step.)

[0347] Step 5:

[0348] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video.

[0349] Input: Search term if no matching sign language word was found.

[0350] Data processing: Input a prompt message into a generative artificial intelligence (e.g., "Please generate a sign language video of 'hello'").

[0351] Output: Generated sign language video

[0352] Step 6:

[0353] The server saves the generated sign language videos as a cache in the database.

[0354] Input: Generated sign language video

[0355] Data processing: Caching of sign language videos and their metadata

[0356] Output: Sign language videos cached in the database

[0357] Step 7:

[0358] The server constructs response data containing the URL of the sign language video and sends it to the terminal in JSON format.

[0359] Input: URL of the sign language video (Example: "http: / / example.com / signs / こんにちは.mp4")

[0360] Data processing: Construction of response data (Example: {"word": "Hello", "video_url": "http: / / example.com / signs / Hello.mp4"})

[0361] Output: Response data is sent to the terminal.

[0362] Step 8:

[0363] The terminal processes the response received from the server, extracts the URL of the sign language video, and initializes the video player.

[0364] Input: Response data (Example: {"word": "Hello", "video_url": "http: / / example.com / signs / Hello.mp4"})

[0365] Data processing: URL extraction and video player initialization.

[0366] Output: Start streaming the sign language video in the video player.

[0367] Step 9:

[0368] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements.

[0369] Input: Sign language video (Example: Sign language video for "hello")

[0370] Action: Practice sign language while watching a video.

[0371] Output: User's proficiency in sign language improves.

[0372] Step 10:

[0373] The device captures the user's facial expressions via its camera, and an emotion engine analyzes those expressions to recognize the user's emotions.

[0374] Input: User's facial expression

[0375] Data processing: Analysis of facial expression data using an emotion engine.

[0376] Output: User sentiment data

[0377] Step 11:

[0378] Based on the user's emotions recognized by the emotion engine, the device adjusts its learning content.

[0379] Input: User sentiment data

[0380] Data processing: Adjusting learning content based on emotional data (e.g., providing supplementary materials)

[0381] Output: Adjusted learning content is provided.

[0382] Step 12:

[0383] The device recommends appropriate sign language learning content based on the user's emotional state.

[0384] Input: User sentiment data

[0385] Data processing: Evaluation of training data and determination of recommendations.

[0386] Output: Recommended sign language learning content

[0387] Step 13:

[0388] After learning sign language, the device automatically displays information about volunteer activities to the user.

[0389] Input: Learning completed

[0390] Data processing: Generating volunteer activity guides

[0391] Output: Information about volunteer activities will be displayed.

[0392] Step 14:

[0393] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0394] Input: Volunteer activity information

[0395] Operation: User registration

[0396] Output: Participation in volunteer activity is complete.

[0397] (Application Example 2)

[0398] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0399] In modern brick-and-mortar stores, there is a lack of support for hearing-impaired individuals to shop smoothly, particularly in the provision of product information and guidance in sign language. Furthermore, there are insufficient means of providing emotionally appropriate learning content to sign language learners. As a result, hearing-impaired individuals and sign language learners are often dissatisfied with their shopping experiences in physical stores.

[0400] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in the database, means for the server to call a generative artificial intelligence and generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to recognize the user's emotions and adjust the learning content, means for the terminal to provide guidance and product information in sign language within a physical store, and means for the terminal to display information about volunteer activities. This makes it possible for people with hearing impairments to smoothly obtain information and shop in physical stores, and for sign language learners to effectively advance their learning.

[0401] "Terminal" refers to a device operated by a user, and includes smartphones, smart glasses, head-mounted displays, or robots.

[0402] A "search box" is an input field where users enter specific words or phrases.

[0403] A "server" is a computer system that receives search requests and performs processing such as accessing databases and using generative artificial intelligence.

[0404] "Sign language words" are specific hand shapes and movements used by people with hearing impairments to communicate.

[0405] "Generative artificial intelligence" is an artificial intelligence technology that generates new sign language videos based on specific input data.

[0406] A "sign language video" is a video that visually displays sign language words and is used to help people with hearing impairments understand them.

[0407] "Emotion recognition" is a technology that captures a user's facial expressions and movements through cameras and sensors and analyzes their emotional state.

[0408] "Adjusting learning content" refers to dynamically changing the difficulty level and content of the learning materials and sign language videos provided according to the user's emotional state.

[0409] A "physical store" is a commercial facility that exists physically, a place where customers can visit in person and purchase goods.

[0410] "Volunteer activities" refer to support activities using sign language and other social contribution activities, and are acts in which users contribute to society through their participation.

[0411] "Response" refers to response data such as search results and sign language videos sent from the server to the terminal.

[0412] This invention relates to a system designed to support the hearing impaired and sign language learners, and is particularly intended to provide users with sign language guidance and product information in physical stores. The system also has a function to recognize the user's emotions and adjust the learning content accordingly.

[0413] Hardware and software to be used

[0414] Hardware for carrying out the present invention includes smart glasses, a camera, and a microphone. As a specific example, the smart glasses are devices such as Google® Glass®. Software includes an emotion recognition engine (e.g., Microsoft® Azure® Face API), a sign language generation AI model, a cloud database (e.g., GOOGLE FI® rebase), and a real-time translation API (e.g., Azure Translator).

[0415] System Configuration

[0416] 1. Terminal Operation: The user wears smart glasses and enters sign language words or instructions into the search box. Input can be done via voice recognition, touchpad, or gestures.

[0417] 2. Sending a search request: The device generates a search request, converts it to JSON format, and sends it to the server. For example, the query "Where is the water?" is sent.

[0418] 3. Server Processing: The server receives the request and extracts the entered word. It then searches the database for the corresponding sign language word. If no search result is found, it invokes a generative artificial intelligence to generate a new sign language video and saves it as a cache.

[0419] 4. Sending the response: The server constructs response data that includes the search results and sends it back to the terminal. For example, it may include a URL for a sign language video of "water".

[0420] 5. Processing the response and playing the video: The terminal processes the received response, extracts the URL of the sign language video, and streams it.

[0421] 6. Emotion Recognition and Learning Content Adjustment: The device captures the user's facial expressions through the camera and analyzes them with an emotion recognition engine. For example, if the system recognizes that the user is confused, it will provide supplementary explanations or additional learning materials.

[0422] 7. In-store guidance: Users can view sign language videos through smart glasses while receiving guidance and product information within a physical store. This allows people with hearing impairments to shop more smoothly.

[0423] 8. Volunteer Activity Information: Finally, the device will display information about volunteer activities after the sign language learning session. If the user is interested, detailed information will be provided, and they can easily register to participate.

[0424] Specific example

[0425] Examples of prompts to input into a generative AI model:

[0426] Please create a video of the sign language word for "water." Clearly demonstrate the hand shape and movement used.

[0427] Thus, this system provides effective support for the hearing impaired and sign language learners, significantly improving the in-store experience. By dynamically adjusting learning content according to the user's emotions, it can provide a more personalized learning experience.

[0428] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0429] Step 1:

[0430] The user wears smart glasses and enters a search query into the search box. Input methods include voice recognition, touchpad, or gesture input. This query may refer to product information or guidance the user wants to know. For example, the user might enter the query, "Where is the water?"

[0431] Input: Query (Example: "Where is the water?")

[0432] Output: Text data of the search query

[0433] Step 2:

[0434] The terminal receives the search query and converts it into a JSON-formatted search request. For example, the query "Where is the water?" is converted to the format {"query": "Where is the water?"}.

[0435] Input: Query text data

[0436] Output: Search request in JSON format

[0437] Step 3:

[0438] The terminal sends the generated search request in JSON format to the server. The data is sent to the server using the HTTP POST method.

[0439] Input: Search request in JSON format

[0440] Output: Request sent to the server

[0441] Step 4:

[0442] The server receives the HTTP request and extracts the search query (word) from the request body. In this example, the word "water" is extracted.

[0443] Input: Request sent to the server

[0444] Output: Extracted word (e.g., "water")

[0445] Step 5:

[0446] The server uses the extracted word to search for the corresponding sign language word in the database. The SQL query SELECT FROM sign_language WHERE word='water'; is executed to retrieve the data corresponding to the sign language word "water" from the database.

[0447] Input: Extracted word

[0448] Output: Search result (data of sign language word)

[0449] Step 6:

[0450] If the server does not find data corresponding to the sign language word in the database, it calls the generative artificial intelligence to generate a new sign language video. The generated sign language video is saved as a cache in the database for future searches.

[0451] Input: Sign language word data (if not present)

[0452] Output: Generated sign language video data

[0453] Step 7:

[0454] The server constructs response data containing the search results or the URL of the generated sign language video and returns it to the terminal. For example, data such as {"word": "water", "video_url": "http: / / example.com / signs / water.mp4"} is returned.

[0455] Input: Search results or data of the generated sign language video

[0456] Output: Response data

[0457] Step 8:

[0458] The terminal processes the response data received from the server and extracts the URL of the sign language video. Then, it initializes the video player on the terminal and streams and plays the sign language video.

[0459] Input: Response data

[0460] Output: Streaming playback of the sign language video

[0461] Step 9:

[0462] The user watches the sign language video played on the terminal and imitates the hand movements to learn sign language. The user practices the sign language for "water" while watching the video.

[0463] Input: Streaming playback of sign language videos

[0464] Output: Learning sign language

[0465] Step 10:

[0466] The device captures the user's facial expressions through its camera and analyzes the user's emotions using an emotion recognition engine (Microsoft Azure Face API). For example, if it captures a confused expression from the user, it analyzes this information.

[0467] Input: Captured user facial expression

[0468] Output: User's emotional state

[0469] Step 11:

[0470] The device adjusts the learning content based on the user's emotional state. If the user is confused, it provides supplementary explanations and additional materials; if the user is in a positive emotional state, it suggests more difficult sign language.

[0471] Input: User's emotional state

[0472] Output: Adjusted learning content

[0473] Step 12:

[0474] After the user finishes learning sign language, the device will display information about volunteer activities. For example, a message such as "Would you like to participate in volunteer activities?" will be displayed.

[0475] Input: Completion of sign language learning

[0476] Output: Display of volunteer activity information

[0477] The above outlines the specific processing steps of how the system of this invention actually works. This system will enable people with hearing impairments and sign language learners to have a better experience in physical stores.

[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0479] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0480] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0481] [Second Embodiment]

[0482] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0483] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0484] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0486] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0488] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0489] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0490] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0491] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0492] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0493] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0494] This invention provides a system that makes it easier for users to learn sign language and also offers tools to promote participation in volunteer activities. The system searches for sign language words entered through a search box and supports learning through images and videos. The system's program processing flow is described below in natural language, with detailed examples.

[0495] System processing flow

[0496] 1. User search for sign language words:

[0497] The user opens the application on their device and enters the sign language word they want to learn into the search box. For example, they might type "hello" and click the search button.

[0498] 2. Sending request data via the terminal:

[0499] The terminal converts the search request containing the words entered by the user into JSON format and sends it to the server using the HTTP POST method. For example, the request body {"query": "Hello"} is generated and sent to the server.

[0500] 3. Receiving and parsing requests by the server:

[0501] The server receives the HTTP request and extracts the search word from the request body. Based on the extracted word "hello," a query is generated to search for the corresponding sign language word in the database.

[0502] 4. Database search by server:

[0503] The server searches the database for the relevant sign language word. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If it does not exist, a generative artificial intelligence is called to generate a sign language video.

[0504] 5. Server-based invocation of generative artificial intelligence:

[0505] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video corresponding to "hello." This generated video is stored in the database as a cache for future searches.

[0506] 6. Server sends response data:

[0507] The server constructs the search result response data and sends it back to the terminal. This response data includes the URL of a sign language video for "hello."

[0508] 7. Processing and display of response data by the terminal:

[0509] The terminal receives the response sent back from the server and extracts the URL of the sign language video. Based on this, it initializes the video player and streams the sign language video to the user.

[0510] 8. User learning of sign language:

[0511] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. For example, a user might practice the sign for "hello" while watching a video.

[0512] 9. Volunteer activity guidance via terminal:

[0513] After watching the video, the device will automatically display information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" will appear on the device.

[0514] 10. User registration for volunteer activities:

[0515] Users follow the instructions and register to participate in volunteer activities that interest them. For example, a user can participate in an activity by filling in the required information on a form and clicking the registration button.

[0516] Thus, the system of the present invention efficiently supports sign language learning and motivates users to participate in volunteer activities. As a result, it can help compensate for the shortage of sign language interpreters and contribute to the promotion of social contribution activities.

[0517] The following describes the processing flow.

[0518] Step 1:

[0519] The user enters the sign language word they want to learn into the search box via their device. After entering the word, they click the search button. For example, they might enter "hello" and press the search button.

[0520] Step 2:

[0521] The terminal receives user input data and generates a search request in JSON format. The generated request is sent to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[0522] Step 3:

[0523] The server receives an HTTP request. It extracts the search term "hello" from the received request body and generates the query necessary for the database search.

[0524] Step 4:

[0525] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[0526] Step 5:

[0527] The server checks the search results from the database. If a matching sign language video exists in the database, it retrieves the URL of that video. If it does not exist, it proceeds to the next step.

[0528] Step 6:

[0529] The server invokes a generative artificial intelligence to generate a sign language video corresponding to the sign language word "hello" requested by the user. This generated video is stored as a cache in the database as feedback.

[0530] Step 7:

[0531] The server constructs response data that includes the URL of the sign language video. For example, it generates data in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[0532] Step 8:

[0533] The server sends the generated response data back to the terminal. The terminal receives the HTTP response.

[0534] Step 9:

[0535] The terminal processes the response received from the server and extracts the URL of the sign language video. Using the extracted URL, the video player is initialized and the sign language video is streamed.

[0536] Step 10:

[0537] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements. Specifically, they practice the sign language for "hello" while watching the video.

[0538] Step 11:

[0539] After watching a sign language video, the device automatically displays information about volunteer activities. For example, a pop-up message might appear asking, "Would you like to participate in volunteer activities?"

[0540] Step 12:

[0541] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0542] (Example 1)

[0543] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0544] Traditional sign language learning systems often lacked relevant video resources when users wanted to learn specific sign language words, which acted as a barrier to learning. Furthermore, they lacked means to motivate participation in volunteer activities in addition to promoting sign language learning. Therefore, there was a need for a system that could simultaneously provide efficient sign language learning support and promote volunteer activities.

[0545] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0546] In this invention, the server includes means for calling a generative artificial intelligence to generate sign language videos, means for storing the sign language videos generated using the generative artificial intelligence as a cache in a database, and means for the server to store search results as a cache in the database. This allows for the real-time generation of sign language videos when a user searches for a specific sign language word, and further enables faster query responses for the same search in the future. In addition, by simultaneously providing information on volunteer activities during the sign language learning process, participation in volunteer activities can also be promoted.

[0547] "User" refers to an individual who uses the sign language learning system.

[0548] A "device" refers to an electronic device used by a user, such as a smartphone, tablet, or personal computer.

[0549] A "search box" refers to the interface used by users to input sign language words they want to learn.

[0550] A "search request" refers to data sent from the device to the server based on the words the user enters into the search box.

[0551] A "server" refers to the central processing unit of a sign language learning system, which is responsible for accessing the database and calling generative artificial intelligence.

[0552] A "database" refers to an information management system that stores sign language words and their associated video data.

[0553] "Generative artificial intelligence" refers to machine learning models and algorithms used to generate sign language videos corresponding to specific sign language words.

[0554] A "sign language video" refers to a video or image that visually demonstrates specific sign language words.

[0555] "Cache" refers to a temporary storage area for data such as generated sign language videos, in order to improve the response speed to future search queries.

[0556] "Streaming playback" refers to the real-time playback of sign language videos on the user's device.

[0557] "Volunteer Activity Guide" refers to information provided to users through the sign language learning system to encourage participation in social contribution activities.

[0558] This invention is a system designed to make it easier for users to learn sign language and to further promote participation in volunteer activities. The system has a mechanism that allows users to input specific sign language words through a search box, and then provides a corresponding sign language video. The specific implementation method of the system is described below.

[0559] Hardware and software configuration

[0560] User terminal

[0561] Users access the system using devices such as smartphones, tablets, or personal computers. These devices require an internet connection, and the system is accessed through a browser application or a dedicated application.

[0562] server

[0563] The server is the central component for receiving and processing user requests. The server implements the following software and functions:

[0564] Web server: Software used to receive HTTP requests and return responses.

[0565] JSON parser: A library for analyzing the format of received data.

[0566] Database Management System: A system for storing and searching sign language words and their corresponding sign language videos.

[0567] Generative artificial intelligence models: Machine learning models and algorithms for generating sign language videos in real time.

[0568] Flow of operations

[0569] 1. User search for sign language words

[0570] The user opens the application on their device and enters the sign language word they want to learn into the search box. For example, they might enter "hello" and click the search button. This sends a search request to the system.

[0571] 2. Sending request data by the terminal

[0572] The terminal converts the search request containing the words entered by the user into JSON format and sends it to the server using the HTTP POST method. Specifically, the generated request data is {"query": "こんにちは"}.

[0573] 3. Receiving and parsing requests by the server

[0574] The server receives the HTTP request and extracts the search word from the request body. Based on the extracted word "hello," a query is generated to search for the corresponding sign language word in the database.

[0575] 4. Server-based database searches and calls to generative artificial intelligence.

[0576] The server searches the database for the relevant sign language word. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If no such data exists in the database, the server calls a generative artificial intelligence to generate a sign language video corresponding to "hello." An example of a prompt to the generative AI model is, "Please generate a sign language video for 'hello'."

[0577] 5. Server sends response data

[0578] The server constructs the search result response data and sends it back to the terminal. This response data includes the URL of a sign language video of "hello." For example, it is sent in the form of {"videoURL": "http: / / example.com / videos / kon-nichiwa.mp4"}.

[0579] 6. Processing and display of response data by the terminal.

[0580] The device receives the response sent back from the server and extracts the URL of the sign language video. Using this URL, the device initializes its video player and streams the sign language video to the user.

[0581] 7. User learning of sign language

[0582] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. For example, a user might practice the sign for "hello."

[0583] 8. Display of volunteer activity information on terminals

[0584] After watching the video, the device automatically displays information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" might appear.

[0585] 9. User registration for volunteer activities

[0586] Users follow the instructions and register to participate in volunteer activities that interest them. For example, they can participate in an activity by filling in the required information on a form and clicking the registration button.

[0587] As described above, this system not only efficiently supports sign language learning but also promotes users' participation in volunteer activities. This helps to alleviate the shortage of sign language interpreters and contributes to the promotion of social contribution activities.

[0588] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0589] Step 1:

[0590] The user launches the application on their device and enters the sign language word they want to learn into the search box. For example, the user enters "hello" and clicks the search button. The device retrieves the contents of the search box based on the user's input and generates JSON data {"query": "hello"}.

[0591] Step 2:

[0592] The terminal sends the generated JSON-formatted search request to the server using the HTTP POST method. Specifically, it uses an HTTP client library to send the request data to the specified endpoint on the server.

[0593] Input: Sign language word "hello" entered in the search box.

[0594] Data processing: Convert search requests to JSON

[0595] Output: Search request in JSON format {"query": "Hello"}

[0596] Step 3:

[0597] The server extracts the request body from the received HTTP request, parses the JSON data to extract the search term "hello". Based on this, the server generates an SQL query SELECT FROM SignLanguageVideos WHERE keyword='hello' to search for the corresponding sign language word in the database.

[0598] Input: HTTP request body

[0599] Data processing: JSON parsing and extraction of search terms.

[0600] Output: SQL query SELECT FROM SignLanguageVideos WHERE keyword='こんにちは'

[0601] Step 4:

[0602] The server executes the generated SQL query against the database to search for the corresponding sign language video. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If it does not exist, a generative artificial intelligence is called to generate a sign language video.

[0603] Input: SQL query

[0604] Data Calculation: Executing Database Queries

[0605] Output: URL of the sign language video or status of not found.

[0606] Step 5:

[0607] If the server does not find a corresponding sign language word in the database, it uses a generative artificial intelligence model to generate a sign language video corresponding to "hello." A prompt message, "Please generate a sign language video for 'hello'," is sent to the generative AI model, and the generated video file is received.

[0608] Input: Search term "hello" - Status: Not found

[0609] Data processing: Generate prompt sentences, send them to an AI model, and generate a video.

[0610] Output: Generated sign language video file

[0611] Step 6:

[0612] The generated sign language videos are stored as a cache in the database to enable faster responses to future queries. This eliminates the need to generate videos again when a search is performed.

[0613] Input: Generated sign language video file

[0614] Data processing: Saving video files

[0615] Output: Sign language video files stored in the database

[0616] Step 7:

[0617] The server constructs response data containing the URL of the generated sign language video or a sign language video retrieved from the database, and sends it back to the terminal. For example, it generates a JSON response such as {"videoURL": "http: / / example.com / videos / kon-nichiwa.mp4"}.

[0618] Input: URL of a sign language video or a generated sign language video

[0619] Data processing: Building response data

[0620] Output: Response data in JSON format

[0621] Step 8:

[0622] The device receives the response sent back from the server and extracts the URL of the sign language video from it. Using the extracted URL, the device initializes the video player and streams the sign language video.

[0623] Input: Response data in JSON format

[0624] Data processing: Parsing responses and extracting URLs.

[0625] Output: Initialization of the video player and streaming playback

[0626] Step 9:

[0627] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. Specifically, they practice the sign for "hello" while watching the videos.

[0628] Input: Watching sign language videos

[0629] Data calculation: Imitation of hand movements

[0630] Output: Learning sign language

[0631] Step 10:

[0632] After watching a video, the device automatically displays information about volunteer activities to the user. For example, it might display a pop-up asking, "Would you like to participate in volunteer activities?"

[0633] Input: Video viewing end event

[0634] Data processing: Displaying volunteer guidance messages

[0635] Output: Display of a popup message

[0636] Step 11:

[0637] Users follow the instructions and register to participate in volunteer activities that interest them. Specifically, they participate in an activity by filling in the required information on a form and clicking the registration button.

[0638] Input: User-generated registration information

[0639] Data processing: Form data submission

[0640] Output: Registration for volunteer activities

[0641] Through the steps outlined above, this system efficiently supports sign language learning and promotes users' participation in volunteer activities.

[0642] (Application Example 1)

[0643] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0644] In modern factory and other work environments, communication between hearing-impaired workers and machines or robots remains a challenge. Furthermore, in environments where understanding sign language is insufficient, smooth communication using sign language is difficult, potentially reducing the work efficiency of hearing-impaired workers. Additionally, a lack of motivation among the general public to learn sign language and participate in volunteer activities is also a problem.

[0645] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0646] In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in a database, means for the server to invoke a generative artificial intelligence to generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to display volunteer activity information, means for capturing the user's sign language movements using a camera and analyzing them, means for sending the analyzed sign language movements to the server and recognizing the instruction content, and means for the robot to perform an action based on the recognized instruction content. This enables smooth communication between workers with hearing impairments and robots in the factory, and makes it possible to promote sign language learning and motivate participation in volunteer activities.

[0647] A "search box" is an input field used by users to enter specific words or phrases via their device.

[0648] A "search request" is a request generated by the device based on data entered by the user and sent to the server.

[0649] A "server" is a computer system that receives and processes search requests.

[0650] A "database" is an information system for systematically storing and managing data such as sign language words and their corresponding videos, in a searchable format.

[0651] "Generative artificial intelligence" refers to artificial intelligence algorithms that can generate new sign language videos based on existing data.

[0652] A "sign language video" is a video recording of actions that express specific sign language words or phrases.

[0653] "Streaming playback" is a method of playing video data in real time, allowing users to watch before the data is fully downloaded.

[0654] A "volunteer activity guide" is a notification or message that provides users with information about volunteer activities and encourages them to participate.

[0655] A "camera" is an imaging device used to capture the user's sign language movements.

[0656] "Analysis" is the process of deciphering sign language movements captured by a camera and determining whether they represent specific sign language words or phrases.

[0657] A "robot" is a mechanical device that autonomously performs specific tasks in factories and other settings, operating based on user instructions.

[0658] A "user" is a person who uses this system to learn sign language or to give instructions to the robot through sign language.

[0659] This invention provides a system that makes it easier for users to learn sign language and enables smooth communication with robots using sign language within a factory. This system is primarily implemented using a user terminal, server, database, camera, and robot.

[0660] User terminal functions

[0661] User devices include smartphones, tablets, and personal computers. These devices are equipped with a search box where users enter sign language words or phrases they wish to learn. A search request is generated and sent to the server. The device has the capability to convert the request into JSON format.

[0662] Server Functions

[0663] The server receives a search request from the user and extracts the words from the request. Next, it searches the database for videos of the corresponding sign language words. If no video exists in the database, it calls upon a generative artificial intelligence to generate a sign language video. The generated video is also stored in the database as a cache for future searches. The server returns the search results to the user's terminal and provides the URL of the sign language video.

[0664] Camera functions

[0665] The camera is used to capture the user's sign language movements. The captured video is analyzed by an internal or external sign language recognition algorithm, and the content of the sign language is extracted. The extracted content is then sent back to the server to be used as instructions for the robot's actions.

[0666] Robot functions

[0667] The robots in the factory perform actions based on sign language instructions received from a server. For example, if a user signs "start work," the robot recognizes the sign and begins the specified task.

[0668] Specific Examples and Generative AI Models

[0669] As a concrete example of this system's use, consider a scenario where a user gives the sign language instruction to "stop work." A camera captures the sign language movement, analyzes the data, and sends it to a server. The server analyzes the data and sends the corresponding instruction to the robot. As a result, the robot immediately stops working. Additionally, when a user searches for sign language videos to learn, the generative artificial intelligence generates new sign language videos and saves them to the database.

[0670] Examples of prompt sentences to use for sign language recognition in a generative AI model are as follows:

[0671] "Design and implement a method to recognize specific sign language actions (e.g., 'start work') from short video clips."

[0672] This system will allow hearing-impaired employees to learn sign language and give natural sign language instructions to robots within the factory, which is expected to improve the working environment and increase operational efficiency.

[0673] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0674] Step 1:

[0675] The user enters the sign language word they want to learn into the search box via their device. The entered word is sent to the device as request data. Specifically, the user enters "hello" and presses the search button. In this case, the input is "hello," and it is recognized by the device.

[0676] Step 2:

[0677] The device generates a search request, converts it to JSON format, and sends it to the server. The generated JSON data is in the format {"query": "こんにちは"} and is sent to the server using the HTTP POST method. The input in this case is the word "こんにちは" entered by the user, and the output is data in JSON format.

[0678] Step 3:

[0679] The server receives a search request and extracts a word from the request body. Specifically, it extracts the value "こんにちは" from the "query" field of the received JSON data. The input in this case is JSON data {"query": "こんにちは"}, and the output is the extracted word "こんにちは".

[0680] Step 4:

[0681] The server searches the database for the corresponding sign language word. Specifically, it uses a database query to search for a sign language video corresponding to "hello." The input in this case is the extracted word "hello," and the output is the URL of the corresponding sign language video.

[0682] Step 5:

[0683] If the server does not find a corresponding sign language video in the database, it invokes a generative artificial intelligence to generate one. The generated sign language video is then saved as a cache in the database. The input in this case is the extracted word "hello," and the output is the generated sign language video.

[0684] Step 6:

[0685] The server returns the search results to the user's device. The search results include URLs for sign language videos. Specifically, the URLs are constructed as a JSON response and sent to the user's device. The input in this process is the search results or the URLs of the generated sign language videos, and the output is the JSON response sent to the user's device.

[0686] Step 7:

[0687] The terminal processes the response received from the server, extracts the URL of the sign language video, and starts streaming playback. Specifically, it initializes the video player and plays the sign language video. The input in this process is a JSON response, and the output is the streaming playback of the sign language video.

[0688] Step 8:

[0689] Users learn sign language by watching sign language videos on their devices. Specifically, they imitate hand movements while watching the videos being played. In this process, the input is the sign language video, and the output is the learning of sign language.

[0690] Step 9:

[0691] The device uses its camera to capture the user's sign language movements and analyzes them. Specifically, it analyzes the video data captured by the camera using an internal algorithm to extract the content of the sign language. The input in this process is the captured sign language video, and the output is the extracted sign language content.

[0692] Step 10:

[0693] The terminal analyzes the sign language gestures and sends them to the server, which then recognizes the instructions. The analysis results are converted to JSON format and sent using the HTTP POST method. The input in this process is the extracted sign language content, and the output is the transmission of JSON data to the server.

[0694] Step 11:

[0695] The robot executes actions based on the instructions recognized by the server. Specifically, the robot receives recognition data from the server and performs the corresponding action, such as "start work." In this process, the input is the recognized instruction, and the output is the robot's action.

[0696] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0697] This invention relates to a system for promoting sign language learning and participation in volunteer activities, with the aim of supporting people with hearing impairments. Furthermore, this invention aims to provide users with an optimal learning experience by combining it with an emotion engine that recognizes the user's emotions. The processing flow of the system's program is described below in natural language and illustrated in detail with specific examples.

[0698] System processing flow

[0699] 1. User search for sign language words:

[0700] The user enters the sign language word they want to learn into the search box via their device and clicks the search button. For example, they might enter "hello".

[0701] 2. Sending request data via the terminal:

[0702] The terminal receives user input data, generates a search request in JSON format, and sends it to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[0703] 3. Receiving and parsing requests by the server:

[0704] The server receives the HTTP request, extracts the search term "hello" from the request body, and generates the query necessary for the database search.

[0705] 4. Database search by server:

[0706] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[0707] 5. Server-based invocation of generative artificial intelligence:

[0708] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video. This generated video is stored in the database as a cache for future searches.

[0709] 6. Server sends response data:

[0710] The server constructs response data containing the URL of the sign language video and sends it back to the terminal. For example, it sends data like {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"} to the terminal.

[0711] 7. Processing and display of response data by the terminal:

[0712] The terminal processes the response received from the server, extracts the URL of the sign language video, initializes the video player, and streams the sign language video.

[0713] 8. User learning of sign language:

[0714] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements. For example, they can practice the sign for "hello" while watching a video.

[0715] 9. User emotion recognition by emotion engine:

[0716] The device captures the user's facial expressions through its camera, and an emotion engine analyzes those expressions to recognize the user's emotions. For example, if the user is smiling during the learning process, it will be recognized as a positive emotion.

[0717] 10. Adjusting the learning content:

[0718] Based on the user's emotions recognized by the emotion engine, the device adjusts the learning content. For example, if it detects a confused expression on the user's face, it will provide supplementary learning materials or explanations.

[0719] 11. Recommendation of appropriate sign language learning content:

[0720] The device recommends appropriate sign language learning content based on the user's emotional state. For example, if the user is in a positive emotional state, it will suggest more difficult sign language.

[0721] 12. Display of volunteer activity information on terminals:

[0722] After learning sign language, the device automatically displays information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" might appear.

[0723] 13. User registration for volunteer activities:

[0724] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0725] As described above, the system of the present invention efficiently supports sign language learning and provides a more effective learning experience by adjusting the learning content according to the user's emotions. Furthermore, it can enhance social contribution by promoting participation in volunteer activities.

[0726] The following describes the processing flow.

[0727] Step 1:

[0728] The user enters the sign language word they want to learn into the search box on their device and clicks the search button. For example, they might enter "hello".

[0729] Step 2:

[0730] The terminal receives user input data and generates a search request in JSON format. The generated JSON request is sent to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[0731] Step 3:

[0732] The server receives an HTTP request. After receiving it, it extracts the search term "hello" from the request body and generates the query necessary for the database search. Specifically, it parses the request and identifies the word entered by the user.

[0733] Step 4:

[0734] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[0735] Step 5:

[0736] The server checks the search results from the database. If a matching sign language video exists in the database, it retrieves the URL of that video. If it does not exist, it proceeds to the next step.

[0737] Step 6:

[0738] The server invokes a generative artificial intelligence to generate a sign language video corresponding to the sign language word "hello" requested by the user. This generated video is stored as a cache in the database for future searches.

[0739] Step 7:

[0740] The server constructs response data that includes the URL of the sign language video. For example, it generates data in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[0741] Step 8:

[0742] The server sends the generated response data back to the user's terminal. The terminal receives the HTTP response.

[0743] Step 9:

[0744] The terminal analyzes the response received from the server and extracts the URL of the sign language video. Then, it initializes the video player and prepares it for streaming playback of the sign language video.

[0745] Step 10:

[0746] Users watch sign language videos on their devices and learn sign language by imitating the hand movements. Specifically, they practice the sign language for "hello" while watching the videos.

[0747] Step 11:

[0748] The device uses its camera to capture the user's facial expressions, and an emotion engine analyzes those expressions to recognize the user's emotions. For example, if the user is confused, the system analyzes and understands that expression.

[0749] Step 12:

[0750] Based on the user's emotions recognized by the emotion engine, the device adjusts the learning content. Specifically, if the user is confused, it displays supplementary materials or additional explanations.

[0751] Step 13:

[0752] The device recommends appropriate sign language learning content based on the user's emotional state. For example, if the user is in a positive state, it will suggest more difficult sign language.

[0753] Step 14:

[0754] After watching a sign language video, the device automatically displays information about social activities and volunteer opportunities to the user. For example, a pop-up message might appear asking, "Would you like to participate in volunteer activities?"

[0755] Step 15:

[0756] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0757] (Example 2)

[0758] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0759] Traditional sign language learning systems had problems such as the inability for users to easily obtain videos corresponding to the sign language words they searched for, which reduced the efficiency of sign language learning. Furthermore, the lack of customization of learning content based on the user's emotional state meant that learning effectiveness was not maximized, potentially leading to decreased user motivation. In addition, there was a lack of means to encourage participation in volunteer activities, resulting in fewer opportunities to raise awareness of social contribution.

[0760] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in the database, means for the server to call a generative artificial intelligence and generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to capture the user's facial expression via a camera and for the emotion engine to recognize the user's emotion, means for the terminal to adjust the learning content based on the recognized emotion, and means for the terminal to display volunteer activity information. As a result, the user can efficiently proceed with learning sign language, adjust the learning content based on their emotional state, and further increase their motivation to participate in volunteer activities.

[0761] A "terminal" is a device that a user directly operates to input and display information.

[0762] A "server" is a device that processes data sent from a terminal and provides the necessary information.

[0763] A "search box" is a graphical user interface element that users use to enter specific information.

[0764] A "search request" is a request made to query a server for necessary data based on information entered by the user.

[0765] "JSON format" is an abbreviation for JavaScript Object Notation, and it is a lightweight text data exchange format for structuring and representing data.

[0766] The "HTTP POST method" is an HTTP request method used by a client to send resources to a server.

[0767] An "emotion engine" is software or a device used to recognize and analyze a user's emotions.

[0768] A "sign language video" is a video recording of sign language movements and expressions.

[0769] "Generative artificial intelligence" is an artificial intelligence technology that has the ability to generate new content based on a given prompt.

[0770] A "prompt message" is text data input to a generative artificial intelligence system to generate a specific response.

[0771] "Streaming playback" is a technology that allows playback while receiving data in real time.

[0772] "Volunteer Activity Guide" refers to guide information displayed to users to obtain information about volunteer activities and participate in them.

[0773] Modes for carrying out the invention

[0774] This invention relates to a sign language learning system for supporting people with hearing impairments and a system for promoting participation in volunteer activities. The system aims to provide an optimal learning experience by recognizing the user's emotions. The following describes in detail the embodiments for specifically carrying out this invention.

[0775] The user enters the sign language word they want to learn into the search box via their device and clicks the search button. Upon receiving this input, the device generates a search request in JSON format and sends it to the server using the HTTP POST method. For example, if the user wants to learn the sign language word for "hello," they would search for "hello" in the input, and the device would generate JSON data {"query": "hello"} and send it to the server.

[0776] The server receives this HTTP request and extracts the search word from the request body. Based on the extracted word, the server generates an SQL query to search for the corresponding sign language word in the database and executes the database search. For example, it executes the SQL query SELECT FROM sign_language WHERE word='こんにちは';

[0777] If the database search does not find a matching sign language word, the server calls a generative artificial intelligence (generative AI model) and generates a sign language video based on the prompt text. This generated sign language video is stored in the database as a cache for future searches. For example, the generative AI model will generate a sign language video if the prompt text "Please generate a sign language video for 'hello'" is entered into it.

[0778] A response data containing the URL of the generated or retrieved sign language video is constructed, and the server returns it to the terminal in JSON format. For example, it might be returned in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[0779] The device processes the received response data, extracts the URL of the sign language video, and initializes the video player. It then streams the sign language video, allowing the user to watch and learn from it on their device.

[0780] Furthermore, the device captures the user's facial expressions via the camera, and an emotion engine analyzes this facial data to recognize the user's emotions. For example, if the user is learning sign language with a smile, a positive emotion will be recognized. Based on this emotion data, the device adjusts the learning content. For instance, if the user shows a confused expression, the device will provide additional learning materials or clearer explanations.

[0781] After completing sign language learning, the device automatically displays information about volunteer activities and encourages the user to participate. For example, it might display a pop-up asking, "Would you like to participate in volunteer activities?" The user can then register to participate by following this prompt.

[0782] This system allows users to learn sign language efficiently, adjust learning content according to their emotions, and foster a sense of social responsibility by encouraging participation in volunteer activities.

[0783] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0784] Step 1:

[0785] The user enters the sign language word they want to learn into the search box via their device and clicks the search button.

[0786] Input: Sign language words entered by the user into the search box on the device (e.g., "hello")

[0787] Output: Clicking the search button sends the entered sign language word as a search request to the terminal.

[0788] Step 2:

[0789] The terminal receives user input data, generates a search request in JSON format, and sends it to the server using the HTTP POST method.

[0790] Input: Sign language word (e.g., "hello")

[0791] Data processing: Generate JSON data containing sign language words (e.g., {"query": "Hello"}).

[0792] Output: JSON data is sent to the server using the HTTP POST method.

[0793] Step 3:

[0794] The server receives an HTTP request and extracts the search term from the request body.

[0795] Input: JSON data sent to the server

[0796] Data processing: Extraction of search terms (e.g., "hello")

[0797] Output: Extracts search terms and prepares a query for database search.

[0798] Step 4:

[0799] The server searches the database for the corresponding sign language word.

[0800] Input: Search term (e.g., "hello")

[0801] Data calculation: Executing database queries (e.g., SELECT FROM sign_language WHERE word='こんにちは';)

[0802] Output: Search results (If a relevant sign language video is found, that will be the result; otherwise, proceed to the next step.)

[0803] Step 5:

[0804] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video.

[0805] Input: Search term if no matching sign language word was found.

[0806] Data processing: Input a prompt message into a generative artificial intelligence (e.g., "Please generate a sign language video of 'hello'").

[0807] Output: Generated sign language video

[0808] Step 6:

[0809] The server saves the generated sign language videos as a cache in the database.

[0810] Input: Generated sign language video

[0811] Data processing: Caching of sign language videos and their metadata

[0812] Output: Sign language videos cached in the database

[0813] Step 7:

[0814] The server constructs response data containing the URL of the sign language video and sends it to the terminal in JSON format.

[0815] Input: URL of the sign language video (Example: "http: / / example.com / signs / こんにちは.mp4")

[0816] Data processing: Construction of response data (Example: {"word": "Hello", "video_url": "http: / / example.com / signs / Hello.mp4"})

[0817] Output: Response data is sent to the terminal.

[0818] Step 8:

[0819] The terminal processes the response received from the server, extracts the URL of the sign language video, and initializes the video player.

[0820] Input: Response data (Example: {"word": "Hello", "video_url": "http: / / example.com / signs / Hello.mp4"})

[0821] Data processing: URL extraction and video player initialization.

[0822] Output: Start streaming the sign language video in the video player.

[0823] Step 9:

[0824] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements.

[0825] Input: Sign language video (Example: Sign language video for "hello")

[0826] Action: Practice sign language while watching a video.

[0827] Output: User's proficiency in sign language improves.

[0828] Step 10:

[0829] The device captures the user's facial expressions via its camera, and an emotion engine analyzes those expressions to recognize the user's emotions.

[0830] Input: User's facial expression

[0831] Data processing: Analysis of facial expression data using an emotion engine.

[0832] Output: User sentiment data

[0833] Step 11:

[0834] Based on the user's emotions recognized by the emotion engine, the device adjusts its learning content.

[0835] Input: User sentiment data

[0836] Data processing: Adjusting learning content based on emotional data (e.g., providing supplementary materials)

[0837] Output: Adjusted learning content is provided.

[0838] Step 12:

[0839] The device recommends appropriate sign language learning content based on the user's emotional state.

[0840] Input: User sentiment data

[0841] Data processing: Evaluation of training data and determination of recommendations.

[0842] Output: Recommended sign language learning content

[0843] Step 13:

[0844] After learning sign language, the device automatically displays information about volunteer activities to the user.

[0845] Input: Learning completed

[0846] Data processing: Generating volunteer activity guides

[0847] Output: Information about volunteer activities will be displayed.

[0848] Step 14:

[0849] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0850] Input: Volunteer activity information

[0851] Operation: User registration

[0852] Output: Participation in volunteer activity is complete.

[0853] (Application Example 2)

[0854] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0855] In modern brick-and-mortar stores, there is a lack of support for hearing-impaired individuals to shop smoothly, particularly in the provision of product information and guidance in sign language. Furthermore, there are insufficient means of providing emotionally appropriate learning content to sign language learners. As a result, hearing-impaired individuals and sign language learners are often dissatisfied with their shopping experiences in physical stores.

[0856] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in the database, means for the server to call a generative artificial intelligence and generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to recognize the user's emotions and adjust the learning content, means for the terminal to provide guidance and product information in sign language within a physical store, and means for the terminal to display information about volunteer activities. This makes it possible for people with hearing impairments to smoothly obtain information and shop in physical stores, and for sign language learners to effectively advance their learning.

[0857] "Terminal" refers to a device operated by a user, and includes smartphones, smart glasses, head-mounted displays, or robots.

[0858] A "search box" is an input field where users enter specific words or phrases.

[0859] A "server" is a computer system that receives search requests and performs processing such as accessing databases and using generative artificial intelligence.

[0860] "Sign language words" are specific hand shapes and movements used by people with hearing impairments to communicate.

[0861] "Generative artificial intelligence" is an artificial intelligence technology that generates new sign language videos based on specific input data.

[0862] A "sign language video" is a video that visually displays sign language words and is used to help people with hearing impairments understand them.

[0863] "Emotion recognition" is a technology that captures a user's facial expressions and movements through cameras and sensors and analyzes their emotional state.

[0864] "Adjusting learning content" refers to dynamically changing the difficulty level and content of the learning materials and sign language videos provided according to the user's emotional state.

[0865] A "physical store" is a commercial facility that exists physically, a place where customers can visit in person and purchase goods.

[0866] "Volunteer activities" refer to support activities using sign language and other social contribution activities, and are acts in which users contribute to society through their participation.

[0867] "Response" refers to response data such as search results and sign language videos sent from the server to the terminal.

[0868] This invention relates to a system designed to support the hearing impaired and sign language learners, and is particularly intended to provide users with sign language guidance and product information in physical stores. The system also has a function to recognize the user's emotions and adjust the learning content accordingly.

[0869] Hardware and software to be used

[0870] The hardware for carrying out the present invention includes smart glasses, a camera, and a microphone. As a specific example, the smart glasses are suitable devices such as Google Glass. The software includes an emotion recognition engine (e.g., Microsoft Azure Face API), a sign language generation AI model, a cloud database (e.g., Google Firebase), and a real-time translation API (e.g., Azure Translator).

[0871] System Configuration

[0872] 1. Terminal Operation: The user wears smart glasses and enters sign language words or instructions into the search box. Input can be done via voice recognition, touchpad, or gestures.

[0873] 2. Sending a search request: The device generates a search request, converts it to JSON format, and sends it to the server. For example, the query "Where is the water?" is sent.

[0874] 3. Server Processing: The server receives the request and extracts the entered word. It then searches the database for the corresponding sign language word. If no search result is found, it invokes a generative artificial intelligence to generate a new sign language video and saves it as a cache.

[0875] 4. Sending the response: The server constructs response data that includes the search results and sends it back to the terminal. For example, it may include a URL for a sign language video of "water".

[0876] 5. Processing the response and playing the video: The terminal processes the received response, extracts the URL of the sign language video, and streams it.

[0877] 6. Emotion Recognition and Learning Content Adjustment: The device captures the user's facial expressions through the camera and analyzes them with an emotion recognition engine. For example, if the system recognizes that the user is confused, it will provide supplementary explanations or additional learning materials.

[0878] 7. In-store guidance: Users can view sign language videos through smart glasses while receiving guidance and product information within a physical store. This allows people with hearing impairments to shop more smoothly.

[0879] 8. Volunteer Activity Information: Finally, the device will display information about volunteer activities after the sign language learning session. If the user is interested, detailed information will be provided, and they can easily register to participate.

[0880] Specific example

[0881] Examples of prompts to input into a generative AI model:

[0882] Please create a video of the sign language word for "water." Clearly demonstrate the hand shape and movement used.

[0883] Thus, this system provides effective support for the hearing impaired and sign language learners, significantly improving the in-store experience. By dynamically adjusting learning content according to the user's emotions, it can provide a more personalized learning experience.

[0884] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0885] Step 1:

[0886] The user wears smart glasses and enters a search query into the search box. Input methods include voice recognition, touchpad, or gesture input. This query may refer to product information or guidance the user wants to know. For example, the user might enter the query, "Where is the water?"

[0887] Input: Query (Example: "Where is the water?")

[0888] Output: Text data of the search query

[0889] Step 2:

[0890] The terminal receives the search query and converts it into a JSON-formatted search request. For example, the query "Where is the water?" is converted to the format {"query": "Where is the water?"}.

[0891] Input: Query text data

[0892] Output: Search request in JSON format

[0893] Step 3:

[0894] The terminal sends the generated search request in JSON format to the server. The data is sent to the server using the HTTP POST method.

[0895] Input: Search request in JSON format

[0896] Output: Request sent to the server

[0897] Step 4:

[0898] The server receives the HTTP request and extracts the search query (word) from the request body. In this example, the word "water" is extracted.

[0899] Input: Request sent to the server

[0900] Output: Extracted word (e.g., "water")

[0901] Step 5:

[0902] The server uses the extracted word to search for the corresponding sign language word in the database. The SQL query SELECT FROM sign_language WHERE word='water'; is executed to retrieve the data corresponding to the sign language word "water" from the database.

[0903] Input: Extracted word

[0904] Output: Search result (data of sign language word)

[0905] Step 6:

[0906] If the server does not find data corresponding to the sign language word in the database, it calls a generative artificial intelligence to generate a new sign language video. The generated sign language video is saved as a cache in the database for future searches.

[0907] Input: Data of sign language word (if not present)

[0908] Output: Generated sign language video data

[0909] Step 7:

[0910] The server constructs response data containing the search result or the URL of the generated sign language video and returns it to the terminal. For example, data such as {"word": "water", "video_url": "http: / / example.com / signs / water.mp4"} is returned.

[0911] Input: Search result or data of the generated sign language video

[0912] Output: Response data

[0913] Step 8:

[0914] The terminal processes the response data received from the server and extracts the URL of the sign language video. Then, it initializes the video player on the terminal and streams and plays the sign language video.

[0915] Input: Response data

[0916] Output: Streaming playback of sign language video

[0917] Step 9:

[0918] The user watches the sign language video played on the terminal and imitates the hand movements to learn sign language. The user practices the sign language for "water" while watching the video.

[0919] Input: Streaming playback of sign language videos

[0920] Output: Learning sign language

[0921] Step 10:

[0922] The device captures the user's facial expressions through its camera and analyzes the user's emotions using an emotion recognition engine (Microsoft Azure Face API). For example, if it captures a confused expression from the user, it analyzes this information.

[0923] Input: Captured user facial expression

[0924] Output: User's emotional state

[0925] Step 11:

[0926] The device adjusts the learning content based on the user's emotional state. If the user is confused, it provides supplementary explanations and additional materials; if the user is in a positive emotional state, it suggests more difficult sign language.

[0927] Input: User's emotional state

[0928] Output: Adjusted learning content

[0929] Step 12:

[0930] After the user finishes learning sign language, the device will display information about volunteer activities. For example, a message such as "Would you like to participate in volunteer activities?" will be displayed.

[0931] Input: Sign language learning completed

[0932] Output: Display of volunteer activity information

[0933] The above outlines the specific processing steps of how the system of this invention actually works. This system will enable people with hearing impairments and sign language learners to have a better experience in physical stores.

[0934] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0935] The data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0936] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0937] [Third Embodiment]

[0938] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0939] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0940] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0941] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0942] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0943] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0944] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0945] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0946] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0947] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0948] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0949] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0950] This invention provides a system that makes it easier for users to learn sign language and also offers tools to promote participation in volunteer activities. The system searches for sign language words entered through a search box and supports learning through images and videos. The system's program processing flow is described below in natural language, with detailed examples.

[0951] System processing flow

[0952] 1. User search for sign language words:

[0953] The user opens the application on their device and enters the sign language word they want to learn into the search box. For example, they might type "hello" and click the search button.

[0954] 2. Sending request data via the terminal:

[0955] The terminal converts the search request containing the words entered by the user into JSON format and sends it to the server using the HTTP POST method. For example, the request body {"query": "Hello"} is generated and sent to the server.

[0956] 3. Receiving and parsing requests by the server:

[0957] The server receives the HTTP request and extracts the search word from the request body. Based on the extracted word "hello," a query is generated to search for the corresponding sign language word in the database.

[0958] 4. Database search by server:

[0959] The server searches the database for the relevant sign language word. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If it does not exist, a generative artificial intelligence is called to generate a sign language video.

[0960] 5. Server-based invocation of generative artificial intelligence:

[0961] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video corresponding to "hello." This generated video is stored in the database as a cache for future searches.

[0962] 6. Server sends response data:

[0963] The server constructs the search result response data and sends it back to the terminal. This response data includes the URL of a sign language video for "hello."

[0964] 7. Processing and display of response data by the terminal:

[0965] The terminal receives the response sent back from the server and extracts the URL of the sign language video. Based on this, it initializes the video player and streams the sign language video to the user.

[0966] 8. User learning of sign language:

[0967] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. For example, a user might practice the sign for "hello" while watching a video.

[0968] 9. Volunteer activity guidance via terminal:

[0969] After watching the video, the device will automatically display information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" will appear on the device.

[0970] 10. User registration for volunteer activities:

[0971] Users follow the instructions and register to participate in volunteer activities that interest them. For example, a user can participate in an activity by filling in the required information on a form and clicking the registration button.

[0972] Thus, the system of the present invention efficiently supports sign language learning and motivates users to participate in volunteer activities. As a result, it can help compensate for the shortage of sign language interpreters and contribute to the promotion of social contribution activities.

[0973] The following describes the processing flow.

[0974] Step 1:

[0975] The user enters the sign language word they want to learn into the search box via their device. After entering the word, they click the search button. For example, they might enter "hello" and press the search button.

[0976] Step 2:

[0977] The terminal receives user input data and generates a search request in JSON format. The generated request is sent to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[0978] Step 3:

[0979] The server receives an HTTP request. It extracts the search term "hello" from the received request body and generates the query necessary for the database search.

[0980] Step 4:

[0981] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[0982] Step 5:

[0983] The server checks the search results from the database. If a matching sign language video exists in the database, it retrieves the URL of that video. If it does not exist, it proceeds to the next step.

[0984] Step 6:

[0985] The server invokes a generative artificial intelligence to generate a sign language video corresponding to the sign language word "hello" requested by the user. This generated video is stored as a cache in the database as feedback.

[0986] Step 7:

[0987] The server constructs response data that includes the URL of the sign language video. For example, it generates data in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[0988] Step 8:

[0989] The server sends the generated response data back to the terminal. The terminal receives the HTTP response.

[0990] Step 9:

[0991] The terminal processes the response received from the server and extracts the URL of the sign language video. Using the extracted URL, the video player is initialized and the sign language video is streamed.

[0992] Step 10:

[0993] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements. Specifically, they practice the sign language for "hello" while watching the video.

[0994] Step 11:

[0995] After watching a sign language video, the device automatically displays information about volunteer activities. For example, a pop-up message might appear asking, "Would you like to participate in volunteer activities?"

[0996] Step 12:

[0997] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[0998] (Example 1)

[0999] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1000] Traditional sign language learning systems often lacked relevant video resources when users wanted to learn specific sign language words, which acted as a barrier to learning. Furthermore, they lacked means to motivate participation in volunteer activities in addition to promoting sign language learning. Therefore, there was a need for a system that could simultaneously provide efficient sign language learning support and promote volunteer activities.

[1001] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1002] In this invention, the server includes means for calling a generative artificial intelligence to generate sign language videos, means for storing the sign language videos generated using the generative artificial intelligence as a cache in a database, and means for the server to store search results as a cache in the database. This allows for the real-time generation of sign language videos when a user searches for a specific sign language word, and further enables faster query responses for the same search in the future. In addition, by simultaneously providing information on volunteer activities during the sign language learning process, participation in volunteer activities can also be promoted.

[1003] "User" refers to an individual who uses the sign language learning system.

[1004] A "device" refers to an electronic device used by a user, such as a smartphone, tablet, or personal computer.

[1005] A "search box" refers to the interface used by users to input sign language words they want to learn.

[1006] A "search request" refers to data sent from the device to the server based on the words the user enters into the search box.

[1007] A "server" refers to the central processing unit of a sign language learning system, which is responsible for accessing the database and calling generative artificial intelligence.

[1008] A "database" refers to an information management system that stores sign language words and their associated video data.

[1009] "Generative artificial intelligence" refers to machine learning models and algorithms used to generate sign language videos corresponding to specific sign language words.

[1010] A "sign language video" refers to a video or image that visually demonstrates specific sign language words.

[1011] "Cache" refers to a temporary storage area for data such as generated sign language videos, in order to improve the response speed to future search queries.

[1012] "Streaming playback" refers to the real-time playback of sign language videos on the user's device.

[1013] "Volunteer Activity Guide" refers to information provided to users through the sign language learning system to encourage participation in social contribution activities.

[1014] This invention is a system designed to make it easier for users to learn sign language and to further promote participation in volunteer activities. The system has a mechanism that allows users to input specific sign language words through a search box, and then provides a corresponding sign language video. The specific implementation method of the system is described below.

[1015] Hardware and software configuration

[1016] User terminal

[1017] Users access the system using devices such as smartphones, tablets, or personal computers. These devices require an internet connection, and the system is accessed through a browser application or a dedicated application.

[1018] server

[1019] The server is the central component for receiving and processing user requests. The server implements the following software and functions:

[1020] Web server: Software used to receive HTTP requests and return responses.

[1021] JSON parser: A library for analyzing the format of received data.

[1022] Database Management System: A system for storing and searching sign language words and their corresponding sign language videos.

[1023] Generative artificial intelligence models: Machine learning models and algorithms for generating sign language videos in real time.

[1024] Flow of operations

[1025] 1. User search for sign language words

[1026] The user opens the application on their device and enters the sign language word they want to learn into the search box. For example, they might enter "hello" and click the search button. This sends a search request to the system.

[1027] 2. Sending request data by the terminal

[1028] The terminal converts the search request containing the words entered by the user into JSON format and sends it to the server using the HTTP POST method. Specifically, the generated request data is {"query": "こんにちは"}.

[1029] 3. Receiving and parsing requests by the server

[1030] The server receives the HTTP request and extracts the search word from the request body. Based on the extracted word "hello," a query is generated to search for the corresponding sign language word in the database.

[1031] 4. Server-based database searches and calls to generative artificial intelligence.

[1032] The server searches the database for the relevant sign language word. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If no such data exists in the database, the server calls a generative artificial intelligence to generate a sign language video corresponding to "hello." An example of a prompt to the generative AI model is, "Please generate a sign language video for 'hello'."

[1033] 5. Server sends response data

[1034] The server constructs the search result response data and sends it back to the terminal. This response data includes the URL of a sign language video of "hello." For example, it is sent in the form of {"videoURL": "http: / / example.com / videos / kon-nichiwa.mp4"}.

[1035] 6. Processing and display of response data by the terminal.

[1036] The device receives the response sent back from the server and extracts the URL of the sign language video. Using this URL, the device initializes its video player and streams the sign language video to the user.

[1037] 7. User learning of sign language

[1038] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. For example, a user might practice the sign for "hello."

[1039] 8. Display of volunteer activity information on terminals

[1040] After watching the video, the device automatically displays information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" might appear.

[1041] 9. User registration for volunteer activities

[1042] Users follow the instructions and register to participate in volunteer activities that interest them. For example, they can participate in an activity by filling in the required information on a form and clicking the registration button.

[1043] As described above, this system not only efficiently supports sign language learning but also promotes users' participation in volunteer activities. This helps to alleviate the shortage of sign language interpreters and contributes to the promotion of social contribution activities.

[1044] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1045] Step 1:

[1046] The user launches the application on their device and enters the sign language word they want to learn into the search box. For example, the user enters "hello" and clicks the search button. The device retrieves the contents of the search box based on the user's input and generates JSON data {"query": "hello"}.

[1047] Step 2:

[1048] The terminal sends the generated JSON-formatted search request to the server using the HTTP POST method. Specifically, it uses an HTTP client library to send the request data to the specified endpoint on the server.

[1049] Input: Sign language word "hello" entered in the search box.

[1050] Data processing: Convert search requests to JSON

[1051] Output: Search request in JSON format {"query": "Hello"}

[1052] Step 3:

[1053] The server extracts the request body from the received HTTP request, parses the JSON data to extract the search term "hello". Based on this, the server generates an SQL query SELECT FROM SignLanguageVideos WHERE keyword='hello' to search for the corresponding sign language word in the database.

[1054] Input: HTTP request body

[1055] Data processing: JSON parsing and extraction of search terms.

[1056] Output: SQL query SELECT FROM SignLanguageVideos WHERE keyword='こんにちは'

[1057] Step 4:

[1058] The server executes the generated SQL query against the database to search for the corresponding sign language video. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If it does not exist, a generative artificial intelligence is called to generate a sign language video.

[1059] Input: SQL query

[1060] Data Calculation: Executing Database Queries

[1061] Output: URL of the sign language video or status of not found.

[1062] Step 5:

[1063] If the server does not find a corresponding sign language word in the database, it uses a generative artificial intelligence model to generate a sign language video corresponding to "hello." A prompt message, "Please generate a sign language video for 'hello'," is sent to the generative AI model, and the generated video file is received.

[1064] Input: Search term "hello" - Status: Not found

[1065] Data processing: Generate prompt sentences, send them to an AI model, and generate a video.

[1066] Output: Generated sign language video file

[1067] Step 6:

[1068] The generated sign language videos are stored as a cache in the database to enable faster responses to future queries. This eliminates the need to generate videos again when a search is performed.

[1069] Input: Generated sign language video file

[1070] Data processing: Saving video files

[1071] Output: Sign language video files stored in the database

[1072] Step 7:

[1073] The server constructs response data containing the URL of the generated sign language video or a sign language video retrieved from the database, and sends it back to the terminal. For example, it generates a JSON response such as {"videoURL": "http: / / example.com / videos / kon-nichiwa.mp4"}.

[1074] Input: URL of a sign language video or a generated sign language video

[1075] Data processing: Building response data

[1076] Output: Response data in JSON format

[1077] Step 8:

[1078] The device receives the response sent back from the server and extracts the URL of the sign language video from it. Using the extracted URL, the device initializes the video player and streams the sign language video.

[1079] Input: Response data in JSON format

[1080] Data processing: Parsing responses and extracting URLs.

[1081] Output: Initialization of the video player and streaming playback

[1082] Step 9:

[1083] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. Specifically, they practice the sign for "hello" while watching the videos.

[1084] Input: Watching sign language videos

[1085] Data calculation: Imitation of hand movements

[1086] Output: Learning sign language

[1087] Step 10:

[1088] After watching a video, the device automatically displays information about volunteer activities to the user. For example, it might display a pop-up asking, "Would you like to participate in volunteer activities?"

[1089] Input: Video viewing end event

[1090] Data processing: Displaying volunteer guidance messages

[1091] Output: Display of a popup message

[1092] Step 11:

[1093] Users follow the instructions and register to participate in volunteer activities that interest them. Specifically, they participate in an activity by filling in the required information on a form and clicking the registration button.

[1094] Input: User-generated registration information

[1095] Data processing: Form data submission

[1096] Output: Registration for volunteer activities

[1097] Through the steps outlined above, this system efficiently supports sign language learning and promotes users' participation in volunteer activities.

[1098] (Application Example 1)

[1099] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1100] In modern factory and other work environments, communication between hearing-impaired workers and machines or robots remains a challenge. Furthermore, in environments where understanding sign language is insufficient, smooth communication using sign language is difficult, potentially reducing the work efficiency of hearing-impaired workers. Additionally, a lack of motivation among the general public to learn sign language and participate in volunteer activities is also a problem.

[1101] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1102] In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in a database, means for the server to invoke a generative artificial intelligence to generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to display volunteer activity information, means for capturing the user's sign language movements using a camera and analyzing them, means for sending the analyzed sign language movements to the server and recognizing the instruction content, and means for the robot to perform an action based on the recognized instruction content. This enables smooth communication between workers with hearing impairments and robots in the factory, and makes it possible to promote sign language learning and motivate participation in volunteer activities.

[1103] A "search box" is an input field used by users to enter specific words or phrases via their device.

[1104] A "search request" is a request generated by the device based on data entered by the user and sent to the server.

[1105] A "server" is a computer system that receives and processes search requests.

[1106] A "database" is an information system for systematically storing and managing data such as sign language words and their corresponding videos, in a searchable format.

[1107] "Generative artificial intelligence" refers to artificial intelligence algorithms that can generate new sign language videos based on existing data.

[1108] A "sign language video" is a video recording of actions that express specific sign language words or phrases.

[1109] "Streaming playback" is a method of playing video data in real time, allowing users to watch before the data is fully downloaded.

[1110] A "volunteer activity guide" is a notification or message that provides users with information about volunteer activities and encourages them to participate.

[1111] A "camera" is an imaging device used to capture the user's sign language movements.

[1112] "Analysis" is the process of deciphering sign language movements captured by a camera and determining whether they represent specific sign language words or phrases.

[1113] A "robot" is a mechanical device that autonomously performs specific tasks in factories and other settings, operating based on user instructions.

[1114] A "user" is a person who uses this system to learn sign language or to give instructions to the robot through sign language.

[1115] This invention provides a system that makes it easier for users to learn sign language and enables smooth communication with robots using sign language within a factory. This system is primarily implemented using a user terminal, server, database, camera, and robot.

[1116] User terminal functions

[1117] User devices include smartphones, tablets, and personal computers. These devices are equipped with a search box where users enter sign language words or phrases they wish to learn. A search request is generated and sent to the server. The device has the capability to convert the request into JSON format.

[1118] Server Functions

[1119] The server receives a search request from the user and extracts the words from the request. Next, it searches the database for videos of the corresponding sign language words. If no video exists in the database, it calls upon a generative artificial intelligence to generate a sign language video. The generated video is also stored in the database as a cache for future searches. The server returns the search results to the user's terminal and provides the URL of the sign language video.

[1120] Camera functions

[1121] The camera is used to capture the user's sign language movements. The captured video is analyzed by an internal or external sign language recognition algorithm, and the content of the sign language is extracted. The extracted content is then sent back to the server to be used as instructions for the robot's actions.

[1122] Robot functions

[1123] The robots in the factory perform actions based on sign language instructions received from a server. For example, if a user signs "start work," the robot recognizes the sign and begins the specified task.

[1124] Specific Examples and Generative AI Models

[1125] As a concrete example of this system's use, consider a scenario where a user gives the sign language instruction to "stop work." A camera captures the sign language movement, analyzes the data, and sends it to a server. The server analyzes the data and sends the corresponding instruction to the robot. As a result, the robot immediately stops working. Additionally, when a user searches for sign language videos to learn, the generative artificial intelligence generates new sign language videos and saves them to the database.

[1126] Examples of prompt sentences to use for sign language recognition in a generative AI model are as follows:

[1127] "Design and implement a method to recognize specific sign language actions (e.g., 'start work') from short video clips."

[1128] This system will allow hearing-impaired employees to learn sign language and give natural sign language instructions to robots within the factory, which is expected to improve the working environment and increase operational efficiency.

[1129] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1130] Step 1:

[1131] The user enters the sign language word they want to learn into the search box via their device. The entered word is sent to the device as request data. Specifically, the user enters "hello" and presses the search button. In this case, the input is "hello," and it is recognized by the device.

[1132] Step 2:

[1133] The device generates a search request, converts it to JSON format, and sends it to the server. The generated JSON data is in the format {"query": "こんにちは"} and is sent to the server using the HTTP POST method. The input in this case is the word "こんにちは" entered by the user, and the output is data in JSON format.

[1134] Step 3:

[1135] The server receives a search request and extracts a word from the request body. Specifically, it extracts the value "こんにちは" from the "query" field of the received JSON data. The input in this case is JSON data {"query": "こんにちは"}, and the output is the extracted word "こんにちは".

[1136] Step 4:

[1137] The server searches the database for the corresponding sign language word. Specifically, it uses a database query to search for a sign language video corresponding to "hello." The input in this case is the extracted word "hello," and the output is the URL of the corresponding sign language video.

[1138] Step 5:

[1139] If the server does not find a corresponding sign language video in the database, it invokes a generative artificial intelligence to generate one. The generated sign language video is then saved as a cache in the database. The input in this case is the extracted word "hello," and the output is the generated sign language video.

[1140] Step 6:

[1141] The server returns the search results to the user's device. The search results include URLs for sign language videos. Specifically, the URLs are constructed as a JSON response and sent to the user's device. The input in this process is the search results or the URLs of the generated sign language videos, and the output is the JSON response sent to the user's device.

[1142] Step 7:

[1143] The terminal processes the response received from the server, extracts the URL of the sign language video, and starts streaming playback. Specifically, it initializes the video player and plays the sign language video. The input in this process is a JSON response, and the output is the streaming playback of the sign language video.

[1144] Step 8:

[1145] Users learn sign language by watching sign language videos on their devices. Specifically, they imitate hand movements while watching the videos being played. In this process, the input is the sign language video, and the output is the learning of sign language.

[1146] Step 9:

[1147] The device uses its camera to capture the user's sign language movements and analyzes them. Specifically, it analyzes the video data captured by the camera using an internal algorithm to extract the content of the sign language. The input in this process is the captured sign language video, and the output is the extracted sign language content.

[1148] Step 10:

[1149] The terminal analyzes the sign language gestures and sends them to the server, which then recognizes the instructions. The analysis results are converted to JSON format and sent using the HTTP POST method. The input in this process is the extracted sign language content, and the output is the transmission of JSON data to the server.

[1150] Step 11:

[1151] The robot executes actions based on the instructions recognized by the server. Specifically, the robot receives recognition data from the server and performs the corresponding action, such as "start work." In this process, the input is the recognized instruction, and the output is the robot's action.

[1152] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1153] This invention relates to a system for promoting sign language learning and participation in volunteer activities, with the aim of supporting people with hearing impairments. Furthermore, this invention aims to provide users with an optimal learning experience by combining it with an emotion engine that recognizes the user's emotions. The processing flow of the system's program is described below in natural language and illustrated in detail with specific examples.

[1154] System processing flow

[1155] 1. User search for sign language words:

[1156] The user enters the sign language word they want to learn into the search box via their device and clicks the search button. For example, they might enter "hello".

[1157] 2. Sending request data via the terminal:

[1158] The terminal receives user input data, generates a search request in JSON format, and sends it to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[1159] 3. Receiving and parsing requests by the server:

[1160] The server receives the HTTP request, extracts the search term "hello" from the request body, and generates the query necessary for the database search.

[1161] 4. Database search by server:

[1162] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[1163] 5. Server-based invocation of generative artificial intelligence:

[1164] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video. This generated video is stored in the database as a cache for future searches.

[1165] 6. Server sends response data:

[1166] The server constructs response data containing the URL of the sign language video and sends it back to the terminal. For example, it sends data like {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"} to the terminal.

[1167] 7. Processing and display of response data by the terminal:

[1168] The terminal processes the response received from the server, extracts the URL of the sign language video, initializes the video player, and streams the sign language video.

[1169] 8. User learning of sign language:

[1170] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements. For example, they can practice the sign for "hello" while watching a video.

[1171] 9. User emotion recognition by emotion engine:

[1172] The device captures the user's facial expressions through its camera, and an emotion engine analyzes those expressions to recognize the user's emotions. For example, if the user is smiling during the learning process, it will be recognized as a positive emotion.

[1173] 10. Adjusting the learning content:

[1174] Based on the user's emotions recognized by the emotion engine, the device adjusts the learning content. For example, if it detects a confused expression on the user's face, it will provide supplementary learning materials or explanations.

[1175] 11. Recommendation of appropriate sign language learning content:

[1176] The device recommends appropriate sign language learning content based on the user's emotional state. For example, if the user is in a positive emotional state, it will suggest more difficult sign language.

[1177] 12. Display of volunteer activity information on terminals:

[1178] After learning sign language, the device automatically displays information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" might appear.

[1179] 13. User registration for volunteer activities:

[1180] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[1181] As described above, the system of the present invention efficiently supports sign language learning and provides a more effective learning experience by adjusting the learning content according to the user's emotions. Furthermore, it can enhance social contribution by promoting participation in volunteer activities.

[1182] The following describes the processing flow.

[1183] Step 1:

[1184] The user enters the sign language word they want to learn into the search box on their device and clicks the search button. For example, they might enter "hello".

[1185] Step 2:

[1186] The terminal receives user input data and generates a search request in JSON format. The generated JSON request is sent to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[1187] Step 3:

[1188] The server receives an HTTP request. After receiving it, it extracts the search term "hello" from the request body and generates the query necessary for the database search. Specifically, it parses the request and identifies the word entered by the user.

[1189] Step 4:

[1190] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[1191] Step 5:

[1192] The server checks the search results from the database. If a matching sign language video exists in the database, it retrieves the URL of that video. If it does not exist, it proceeds to the next step.

[1193] Step 6:

[1194] The server invokes a generative artificial intelligence to generate a sign language video corresponding to the sign language word "hello" requested by the user. This generated video is stored as a cache in the database for future searches.

[1195] Step 7:

[1196] The server constructs response data that includes the URL of the sign language video. For example, it generates data in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[1197] Step 8:

[1198] The server sends the generated response data back to the user's terminal. The terminal receives the HTTP response.

[1199] Step 9:

[1200] The terminal analyzes the response received from the server and extracts the URL of the sign language video. Then, it initializes the video player and prepares it for streaming playback of the sign language video.

[1201] Step 10:

[1202] Users watch sign language videos on their devices and learn sign language by imitating the hand movements. Specifically, they practice the sign language for "hello" while watching the videos.

[1203] Step 11:

[1204] The device uses its camera to capture the user's facial expressions, and an emotion engine analyzes those expressions to recognize the user's emotions. For example, if the user is confused, the system analyzes and understands that expression.

[1205] Step 12:

[1206] Based on the user's emotions recognized by the emotion engine, the device adjusts the learning content. Specifically, if the user is confused, it displays supplementary materials or additional explanations.

[1207] Step 13:

[1208] The device recommends appropriate sign language learning content based on the user's emotional state. For example, if the user is in a positive state, it will suggest more difficult sign language.

[1209] Step 14:

[1210] After watching a sign language video, the device automatically displays information about social activities and volunteer opportunities to the user. For example, a pop-up message might appear asking, "Would you like to participate in volunteer activities?"

[1211] Step 15:

[1212] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[1213] (Example 2)

[1214] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1215] Traditional sign language learning systems had problems such as the inability for users to easily obtain videos corresponding to the sign language words they searched for, which reduced the efficiency of sign language learning. Furthermore, the lack of customization of learning content based on the user's emotional state meant that learning effectiveness was not maximized, potentially leading to decreased user motivation. In addition, there was a lack of means to encourage participation in volunteer activities, resulting in fewer opportunities to raise awareness of social contribution.

[1216] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in the database, means for the server to call a generative artificial intelligence and generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to capture the user's facial expression via a camera and for the emotion engine to recognize the user's emotion, means for the terminal to adjust the learning content based on the recognized emotion, and means for the terminal to display volunteer activity information. As a result, the user can efficiently proceed with learning sign language, adjust the learning content based on their emotional state, and further increase their motivation to participate in volunteer activities.

[1217] A "terminal" is a device that a user directly operates to input and display information.

[1218] A "server" is a device that processes data sent from a terminal and provides the necessary information.

[1219] A "search box" is a graphical user interface element that users use to enter specific information.

[1220] A "search request" is a request made to query a server for necessary data based on information entered by the user.

[1221] "JSON format" is an abbreviation for JavaScript Object Notation, and it is a lightweight text data exchange format for structuring and representing data.

[1222] The "HTTP POST method" is an HTTP request method used by a client to send resources to a server.

[1223] An "emotion engine" is software or a device used to recognize and analyze a user's emotions.

[1224] A "sign language video" is a video recording of sign language movements and expressions.

[1225] "Generative artificial intelligence" is an artificial intelligence technology that has the ability to generate new content based on a given prompt.

[1226] A "prompt message" is text data input to a generative artificial intelligence system to generate a specific response.

[1227] "Streaming playback" is a technology that allows playback while receiving data in real time.

[1228] "Volunteer Activity Guide" refers to guide information displayed to users to obtain information about volunteer activities and participate in them.

[1229] Modes for carrying out the invention

[1230] This invention relates to a sign language learning system for supporting people with hearing impairments and a system for promoting participation in volunteer activities. The system aims to provide an optimal learning experience by recognizing the user's emotions. The following describes in detail the embodiments for specifically carrying out this invention.

[1231] The user enters the sign language word they want to learn into the search box via their device and clicks the search button. Upon receiving this input, the device generates a search request in JSON format and sends it to the server using the HTTP POST method. For example, if the user wants to learn the sign language word for "hello," they would search for "hello" in the input, and the device would generate JSON data {"query": "hello"} and send it to the server.

[1232] The server receives this HTTP request and extracts the search word from the request body. Based on the extracted word, the server generates an SQL query to search for the corresponding sign language word in the database and executes the database search. For example, it executes the SQL query SELECT FROM sign_language WHERE word='こんにちは';

[1233] If the database search does not find a matching sign language word, the server calls a generative artificial intelligence (generative AI model) and generates a sign language video based on the prompt text. This generated sign language video is stored in the database as a cache for future searches. For example, the generative AI model will generate a sign language video if the prompt text "Please generate a sign language video for 'hello'" is entered into it.

[1234] A response data containing the URL of the generated or retrieved sign language video is constructed, and the server returns it to the terminal in JSON format. For example, it might be returned in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[1235] The device processes the received response data, extracts the URL of the sign language video, and initializes the video player. It then streams the sign language video, allowing the user to watch and learn from it on their device.

[1236] Furthermore, the device captures the user's facial expressions via the camera, and an emotion engine analyzes this facial data to recognize the user's emotions. For example, if the user is learning sign language with a smile, a positive emotion will be recognized. Based on this emotion data, the device adjusts the learning content. For instance, if the user shows a confused expression, the device will provide additional learning materials or clearer explanations.

[1237] After completing sign language learning, the device automatically displays information about volunteer activities and encourages the user to participate. For example, it might display a pop-up asking, "Would you like to participate in volunteer activities?" The user can then register to participate by following this prompt.

[1238] This system allows users to learn sign language efficiently, adjust learning content according to their emotions, and foster a sense of social responsibility by encouraging participation in volunteer activities.

[1239] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1240] Step 1:

[1241] The user enters the sign language word they want to learn into the search box via their device and clicks the search button.

[1242] Input: Sign language words entered by the user into the search box on the device (e.g., "hello")

[1243] Output: Clicking the search button sends the entered sign language word as a search request to the terminal.

[1244] Step 2:

[1245] The terminal receives user input data, generates a search request in JSON format, and sends it to the server using the HTTP POST method.

[1246] Input: Sign language word (e.g., "hello")

[1247] Data processing: Generate JSON data containing sign language words (e.g., {"query": "Hello"}).

[1248] Output: JSON data is sent to the server using the HTTP POST method.

[1249] Step 3:

[1250] The server receives an HTTP request and extracts the search term from the request body.

[1251] Input: JSON data sent to the server

[1252] Data processing: Extraction of search terms (e.g., "hello")

[1253] Output: Extracts search terms and prepares a query for database search.

[1254] Step 4:

[1255] The server searches the database for the corresponding sign language word.

[1256] Input: Search term (e.g., "hello")

[1257] Data calculation: Executing database queries (e.g., SELECT FROM sign_language WHERE word='こんにちは';)

[1258] Output: Search results (If a relevant sign language video is found, that will be the result; otherwise, proceed to the next step.)

[1259] Step 5:

[1260] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video.

[1261] Input: Search term if no matching sign language word was found.

[1262] Data processing: Input a prompt message into a generative artificial intelligence (e.g., "Please generate a sign language video of 'hello'").

[1263] Output: Generated sign language video

[1264] Step 6:

[1265] The server saves the generated sign language videos as a cache in the database.

[1266] Input: Generated sign language video

[1267] Data processing: Caching of sign language videos and their metadata

[1268] Output: Sign language videos cached in the database

[1269] Step 7:

[1270] The server constructs response data containing the URL of the sign language video and sends it to the terminal in JSON format.

[1271] Input: URL of the sign language video (Example: "http: / / example.com / signs / こんにちは.mp4")

[1272] Data processing: Construction of response data (Example: {"word": "Hello", "video_url": "http: / / example.com / signs / Hello.mp4"})

[1273] Output: Response data is sent to the terminal.

[1274] Step 8:

[1275] The terminal processes the response received from the server, extracts the URL of the sign language video, and initializes the video player.

[1276] Input: Response data (Example: {"word": "Hello", "video_url": "http: / / example.com / signs / Hello.mp4"})

[1277] Data processing: URL extraction and video player initialization.

[1278] Output: Start streaming the sign language video in the video player.

[1279] Step 9:

[1280] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements.

[1281] Input: Sign language video (Example: Sign language video for "hello")

[1282] Action: Practice sign language while watching a video.

[1283] Output: User's proficiency in sign language improves.

[1284] Step 10:

[1285] The device captures the user's facial expressions via its camera, and an emotion engine analyzes those expressions to recognize the user's emotions.

[1286] Input: User's facial expression

[1287] Data processing: Analysis of facial expression data using an emotion engine.

[1288] Output: User sentiment data

[1289] Step 11:

[1290] Based on the user's emotions recognized by the emotion engine, the device adjusts its learning content.

[1291] Input: User sentiment data

[1292] Data processing: Adjusting learning content based on emotional data (e.g., providing supplementary materials)

[1293] Output: Adjusted learning content is provided.

[1294] Step 12:

[1295] The device recommends appropriate sign language learning content based on the user's emotional state.

[1296] Input: User sentiment data

[1297] Data processing: Evaluation of training data and determination of recommendations.

[1298] Output: Recommended sign language learning content

[1299] Step 13:

[1300] After learning sign language, the device automatically displays information about volunteer activities to the user.

[1301] Input: Learning completed

[1302] Data processing: Generating volunteer activity guides

[1303] Output: Information about volunteer activities will be displayed.

[1304] Step 14:

[1305] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[1306] Input: Volunteer activity information

[1307] Operation: User registration

[1308] Output: Participation in volunteer activity is complete.

[1309] (Application Example 2)

[1310] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1311] In modern brick-and-mortar stores, there is a lack of support for hearing-impaired individuals to shop smoothly, particularly in the provision of product information and guidance in sign language. Furthermore, there are insufficient means of providing emotionally appropriate learning content to sign language learners. As a result, hearing-impaired individuals and sign language learners are often dissatisfied with their shopping experiences in physical stores.

[1312] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in the database, means for the server to call a generative artificial intelligence and generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to recognize the user's emotions and adjust the learning content, means for the terminal to provide guidance and product information in sign language within a physical store, and means for the terminal to display information about volunteer activities. This makes it possible for people with hearing impairments to smoothly obtain information and shop in physical stores, and for sign language learners to effectively advance their learning.

[1313] "Terminal" refers to a device operated by a user, and includes smartphones, smart glasses, head-mounted displays, or robots.

[1314] A "search box" is an input field where users enter specific words or phrases.

[1315] A "server" is a computer system that receives search requests and performs processing such as accessing databases and using generative artificial intelligence.

[1316] "Sign language words" are specific hand shapes and movements used by people with hearing impairments to communicate.

[1317] "Generative artificial intelligence" is an artificial intelligence technology that generates new sign language videos based on specific input data.

[1318] A "sign language video" is a video that visually displays sign language words and is used to help people with hearing impairments understand them.

[1319] "Emotion recognition" is a technology that captures a user's facial expressions and movements through cameras and sensors and analyzes their emotional state.

[1320] "Adjusting learning content" refers to dynamically changing the difficulty level and content of the learning materials and sign language videos provided according to the user's emotional state.

[1321] A "physical store" is a commercial facility that exists physically, a place where customers can visit in person and purchase goods.

[1322] "Volunteer activities" refer to support activities using sign language and other social contribution activities, and are acts in which users contribute to society through their participation.

[1323] "Response" refers to response data such as search results and sign language videos sent from the server to the terminal.

[1324] This invention relates to a system designed to support the hearing impaired and sign language learners, and is particularly intended to provide users with sign language guidance and product information in physical stores. The system also has a function to recognize the user's emotions and adjust the learning content accordingly.

[1325] Hardware and software to be used

[1326] The hardware for carrying out the present invention includes smart glasses, a camera, and a microphone. As a specific example, the smart glasses are suitable devices such as Google Glass. The software includes an emotion recognition engine (e.g., Microsoft Azure Face API), a sign language generation AI model, a cloud database (e.g., Google Firebase), and a real-time translation API (e.g., Azure Translator).

[1327] System Configuration

[1328] 1. Terminal Operation: The user wears smart glasses and enters sign language words or instructions into the search box. Input can be done via voice recognition, touchpad, or gestures.

[1329] 2. Sending a search request: The device generates a search request, converts it to JSON format, and sends it to the server. For example, the query "Where is the water?" is sent.

[1330] 3. Server Processing: The server receives the request and extracts the entered word. It then searches the database for the corresponding sign language word. If no search result is found, it invokes a generative artificial intelligence to generate a new sign language video and saves it as a cache.

[1331] 4. Sending the response: The server constructs response data that includes the search results and sends it back to the terminal. For example, it may include a URL for a sign language video of "water".

[1332] 5. Processing the response and playing the video: The terminal processes the received response, extracts the URL of the sign language video, and streams it.

[1333] 6. Emotion Recognition and Learning Content Adjustment: The device captures the user's facial expressions through the camera and analyzes them with an emotion recognition engine. For example, if the system recognizes that the user is confused, it will provide supplementary explanations or additional learning materials.

[1334] 7. In-store guidance: Users can view sign language videos through smart glasses while receiving guidance and product information within a physical store. This allows people with hearing impairments to shop more smoothly.

[1335] 8. Volunteer Activity Information: Finally, the device will display information about volunteer activities after the sign language learning session. If the user is interested, detailed information will be provided, and they can easily register to participate.

[1336] Specific example

[1337] Examples of prompts to input into a generative AI model:

[1338] Please create a video of the sign language word for "water." Clearly demonstrate the hand shape and movement used.

[1339] Thus, this system provides effective support for the hearing impaired and sign language learners, significantly improving the in-store experience. By dynamically adjusting learning content according to the user's emotions, it can provide a more personalized learning experience.

[1340] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1341] Step 1:

[1342] The user wears smart glasses and enters a search query into the search box. Input methods include voice recognition, touchpad, or gesture input. This query may refer to product information or guidance the user wants to know. For example, the user might enter the query, "Where is the water?"

[1343] Input: Query (Example: "Where is the water?")

[1344] Output: Text data of the search query

[1345] Step 2:

[1346] The terminal receives the search query and converts it into a JSON-formatted search request. For example, the query "Where is the water?" is converted to the format {"query": "Where is the water?"}.

[1347] Input: Query text data

[1348] Output: Search request in JSON format

[1349] Step 3:

[1350] The terminal sends the generated search request in JSON format to the server. The data is sent to the server using the HTTP POST method.

[1351] Input: Search request in JSON format

[1352] Output: Request sent to the server

[1353] Step 4:

[1354] The server receives the HTTP request and extracts the search query (word) from the request body. In this example, the word "water" is extracted.

[1355] Input: Request sent to the server

[1356] Output: Extracted word (e.g., "water")[[]END]

[1357] Step 5:[[ID=3۷]]

[1358] The server uses the extracted word to search for the corresponding sign language word in the database. The SQL query SELECT FROM sign_language WHERE word='water'; is executed to retrieve the data corresponding to the sign language word "water" from the database.

[1359] Input: Extracted word

[1360] Output: Search result (data of sign language word)

[1361] Step 6:

[1362] If the server does not find data corresponding to the sign language word in the database, it calls a generative artificial intelligence to generate a new sign language video. The generated sign language video is saved as a cache in the database for future searches.

[1363] Input: Sign language word data (if not present)

[1364] Output: Generated sign language video data

[1365] Step 7:

[1366] The server constructs response data including the search results or the URL of the generated sign language video and returns it to the terminal. For example, data like {"word": "water", "video_url": "http: / / example.com / signs / water.mp4"} is returned.

[1367] Input: Search results or data of the generated sign language video

[1368] Output: Response data

[1369] Step 8:

[1370] The terminal processes the response data received from the server and extracts the URL of the sign language video. Then, it initializes the video player on the terminal and streams and plays the sign language video.

[1371] Input: Response data

[1372] Output: Streaming playback of the sign language video

[1373] Step 9:

[1374] The user watches the sign language video played on the terminal and imitates the hand movements to learn sign language. The user practices the sign language for "water" while watching the video.

[1375] Input: Streaming playback of sign language videos

[1376] Output: Learning sign language

[1377] Step 10:

[1378] The device captures the user's facial expressions through its camera and analyzes the user's emotions using an emotion recognition engine (Microsoft Azure Face API). For example, if it captures a confused expression from the user, it analyzes this information.

[1379] Input: Captured user facial expression

[1380] Output: User's emotional state

[1381] Step 11:

[1382] The device adjusts the learning content based on the user's emotional state. If the user is confused, it provides supplementary explanations and additional materials; if the user is in a positive emotional state, it suggests more difficult sign language.

[1383] Input: User's emotional state

[1384] Output: Adjusted learning content

[1385] Step 12:

[1386] After the user finishes learning sign language, the device will display information about volunteer activities. For example, a message such as "Would you like to participate in volunteer activities?" will be displayed.

[1387] Input: Sign language learning completed

[1388] Output: Display of volunteer activity information

[1389] The above outlines the specific processing steps of how the system of this invention actually works. This system will enable people with hearing impairments and sign language learners to have a better experience in physical stores.

[1390] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1391] The data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1392] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1393] [Fourth Embodiment]

[1394] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1395] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1396] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1397] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1398] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1399] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1400] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1401] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1402] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1403] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1404] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1405] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1406] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1407] This invention provides a system that makes it easier for users to learn sign language and also offers tools to promote participation in volunteer activities. The system searches for sign language words entered through a search box and supports learning through images and videos. The system's program processing flow is described below in natural language, with detailed examples.

[1408] System processing flow

[1409] 1. User search for sign language words:

[1410] The user opens the application on their device and enters the sign language word they want to learn into the search box. For example, they might type "hello" and click the search button.

[1411] 2. Sending request data via the terminal:

[1412] The terminal converts the search request containing the words entered by the user into JSON format and sends it to the server using the HTTP POST method. For example, the request body {"query": "Hello"} is generated and sent to the server.

[1413] 3. Receiving and parsing requests by the server:

[1414] The server receives the HTTP request and extracts the search word from the request body. Based on the extracted word "hello," a query is generated to search for the corresponding sign language word in the database.

[1415] 4. Database search by server:

[1416] The server searches the database for the relevant sign language word. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If it does not exist, a generative artificial intelligence is called to generate a sign language video.

[1417] 5. Server-based invocation of generative artificial intelligence:

[1418] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video corresponding to "hello." This generated video is stored in the database as a cache for future searches.

[1419] 6. Server sends response data:

[1420] The server constructs the search result response data and sends it back to the terminal. This response data includes the URL of a sign language video for "hello."

[1421] 7. Processing and display of response data by the terminal:

[1422] The terminal receives the response sent back from the server and extracts the URL of the sign language video. Based on this, it initializes the video player and streams the sign language video to the user.

[1423] 8. User learning of sign language:

[1424] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. For example, a user might practice the sign for "hello" while watching a video.

[1425] 9. Volunteer activity guidance via terminal:

[1426] After watching the video, the device will automatically display information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" will appear on the device.

[1427] 10. User registration for volunteer activities:

[1428] Users follow the instructions and register to participate in volunteer activities that interest them. For example, a user can participate in an activity by filling in the required information on a form and clicking the registration button.

[1429] Thus, the system of the present invention efficiently supports sign language learning and motivates users to participate in volunteer activities. As a result, it can help compensate for the shortage of sign language interpreters and contribute to the promotion of social contribution activities.

[1430] The following describes the processing flow.

[1431] Step 1:

[1432] The user enters the sign language word they want to learn into the search box via their device. After entering the word, they click the search button. For example, they might enter "hello" and press the search button.

[1433] Step 2:

[1434] The terminal receives user input data and generates a search request in JSON format. The generated request is sent to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[1435] Step 3:

[1436] The server receives an HTTP request. It extracts the search term "hello" from the received request body and generates the query necessary for the database search.

[1437] Step 4:

[1438] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[1439] Step 5:

[1440] The server checks the search results from the database. If a matching sign language video exists in the database, it retrieves the URL of that video. If it does not exist, it proceeds to the next step.

[1441] Step 6:

[1442] The server invokes a generative artificial intelligence to generate a sign language video corresponding to the sign language word "hello" requested by the user. This generated video is stored as a cache in the database as feedback.

[1443] Step 7:

[1444] The server constructs response data that includes the URL of the sign language video. For example, it generates data in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[1445] Step 8:

[1446] The server sends the generated response data back to the terminal. The terminal receives the HTTP response.

[1447] Step 9:

[1448] The terminal processes the response received from the server and extracts the URL of the sign language video. Using the extracted URL, the video player is initialized and the sign language video is streamed.

[1449] Step 10:

[1450] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements. Specifically, they practice the sign language for "hello" while watching the video.

[1451] Step 11:

[1452] After watching a sign language video, the device automatically displays information about volunteer activities. For example, a pop-up message might appear asking, "Would you like to participate in volunteer activities?"

[1453] Step 12:

[1454] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[1455] (Example 1)

[1456] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1457] Traditional sign language learning systems often lacked relevant video resources when users wanted to learn specific sign language words, which acted as a barrier to learning. Furthermore, they lacked means to motivate participation in volunteer activities in addition to promoting sign language learning. Therefore, there was a need for a system that could simultaneously provide efficient sign language learning support and promote volunteer activities.

[1458] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1459] In this invention, the server includes means for calling a generative artificial intelligence to generate sign language videos, means for storing the sign language videos generated using the generative artificial intelligence as a cache in a database, and means for the server to store search results as a cache in the database. This allows for the real-time generation of sign language videos when a user searches for a specific sign language word, and further enables faster query responses for the same search in the future. In addition, by simultaneously providing information on volunteer activities during the sign language learning process, participation in volunteer activities can also be promoted.

[1460] "User" refers to an individual who uses the sign language learning system.

[1461] A "device" refers to an electronic device used by a user, such as a smartphone, tablet, or personal computer.

[1462] A "search box" refers to the interface used by users to input sign language words they want to learn.

[1463] A "search request" refers to data sent from the device to the server based on the words the user enters into the search box.

[1464] A "server" refers to the central processing unit of a sign language learning system, which is responsible for accessing the database and calling generative artificial intelligence.

[1465] A "database" refers to an information management system that stores sign language words and their associated video data.

[1466] "Generative artificial intelligence" refers to machine learning models and algorithms used to generate sign language videos corresponding to specific sign language words.

[1467] A "sign language video" refers to a video or image that visually demonstrates specific sign language words.

[1468] "Cache" refers to a temporary storage area for data such as generated sign language videos, in order to improve the response speed to future search queries.

[1469] "Streaming playback" refers to the real-time playback of sign language videos on the user's device.

[1470] "Volunteer Activity Guide" refers to information provided to users through the sign language learning system to encourage participation in social contribution activities.

[1471] This invention is a system designed to make it easier for users to learn sign language and to further promote participation in volunteer activities. The system has a mechanism that allows users to input specific sign language words through a search box, and then provides a corresponding sign language video. The specific implementation method of the system is described below.

[1472] Hardware and software configuration

[1473] User terminal

[1474] Users access the system using devices such as smartphones, tablets, or personal computers. These devices require an internet connection, and the system is accessed through a browser application or a dedicated application.

[1475] server

[1476] The server is the central component for receiving and processing user requests. The server implements the following software and functions:

[1477] Web server: Software used to receive HTTP requests and return responses.

[1478] JSON parser: A library for analyzing the format of received data.

[1479] Database Management System: A system for storing and searching sign language words and their corresponding sign language videos.

[1480] Generative artificial intelligence models: Machine learning models and algorithms for generating sign language videos in real time.

[1481] Flow of operations

[1482] 1. User search for sign language words

[1483] The user opens the application on their device and enters the sign language word they want to learn into the search box. For example, they might enter "hello" and click the search button. This sends a search request to the system.

[1484] 2. Sending request data by the terminal

[1485] The terminal converts the search request containing the words entered by the user into JSON format and sends it to the server using the HTTP POST method. Specifically, the generated request data is {"query": "こんにちは"}.

[1486] 3. Receiving and parsing requests by the server

[1487] The server receives the HTTP request and extracts the search word from the request body. Based on the extracted word "hello," a query is generated to search for the corresponding sign language word in the database.

[1488] 4. Server-based database searches and calls to generative artificial intelligence.

[1489] The server searches the database for the relevant sign language word. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If no such data exists in the database, the server calls a generative artificial intelligence to generate a sign language video corresponding to "hello." An example of a prompt to the generative AI model is, "Please generate a sign language video for 'hello'."

[1490] 5. Server sends response data

[1491] The server constructs the search result response data and sends it back to the terminal. This response data includes the URL of a sign language video of "hello." For example, it is sent in the form of {"videoURL": "http: / / example.com / videos / kon-nichiwa.mp4"}.

[1492] 6. Processing and display of response data by the terminal.

[1493] The device receives the response sent back from the server and extracts the URL of the sign language video. Using this URL, the device initializes its video player and streams the sign language video to the user.

[1494] 7. User learning of sign language

[1495] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. For example, a user might practice the sign for "hello."

[1496] 8. Display of volunteer activity information on terminals

[1497] After watching the video, the device automatically displays information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" might appear.

[1498] 9. User registration for volunteer activities

[1499] Users follow the instructions and register to participate in volunteer activities that interest them. For example, they can participate in an activity by filling in the required information on a form and clicking the registration button.

[1500] As described above, this system not only efficiently supports sign language learning but also promotes users' participation in volunteer activities. This helps to alleviate the shortage of sign language interpreters and contributes to the promotion of social contribution activities.

[1501] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1502] Step 1:

[1503] The user launches the application on their device and enters the sign language word they want to learn into the search box. For example, the user enters "hello" and clicks the search button. The device retrieves the contents of the search box based on the user's input and generates JSON data {"query": "hello"}.

[1504] Step 2:

[1505] The terminal sends the generated JSON-formatted search request to the server using the HTTP POST method. Specifically, it uses an HTTP client library to send the request data to the specified endpoint on the server.

[1506] Input: Sign language word "hello" entered in the search box.

[1507] Data processing: Convert search requests to JSON

[1508] Output: Search request in JSON format {"query": "Hello"}

[1509] Step 3:

[1510] The server extracts the request body from the received HTTP request, parses the JSON data to extract the search term "hello". Based on this, the server generates an SQL query SELECT FROM SignLanguageVideos WHERE keyword='hello' to search for the corresponding sign language word in the database.

[1511] Input: HTTP request body

[1512] Data processing: JSON parsing and extraction of search terms.

[1513] Output: SQL query SELECT FROM SignLanguageVideos WHERE keyword='こんにちは'

[1514] Step 4:

[1515] The server executes the generated SQL query against the database to search for the corresponding sign language video. If a sign language video for "hello" exists in the database, its URL is prepared as response data. If it does not exist, a generative artificial intelligence is called to generate a sign language video.

[1516] Input: SQL query

[1517] Data Calculation: Executing Database Queries

[1518] Output: URL of the sign language video or status of not found.

[1519] Step 5:

[1520] If the server does not find a corresponding sign language word in the database, it uses a generative artificial intelligence model to generate a sign language video corresponding to "hello." A prompt message, "Please generate a sign language video for 'hello'," is sent to the generative AI model, and the generated video file is received.

[1521] Input: Search term "hello" - Status: Not found

[1522] Data processing: Generate prompt sentences, send them to an AI model, and generate a video.

[1523] Output: Generated sign language video file

[1524] Step 6:

[1525] The generated sign language videos are stored as a cache in the database to enable faster responses to future queries. This eliminates the need to generate videos again when a search is performed.

[1526] Input: Generated sign language video file

[1527] Data processing: Saving video files

[1528] Output: Sign language video files stored in the database

[1529] Step 7:

[1530] The server constructs response data containing the URL of the generated sign language video or a sign language video retrieved from the database, and sends it back to the terminal. For example, it generates a JSON response such as {"videoURL": "http: / / example.com / videos / kon-nichiwa.mp4"}.

[1531] Input: URL of a sign language video or a generated sign language video

[1532] Data processing: Building response data

[1533] Output: Response data in JSON format

[1534] Step 8:

[1535] The device receives the response sent back from the server and extracts the URL of the sign language video from it. Using the extracted URL, the device initializes the video player and streams the sign language video.

[1536] Input: Response data in JSON format

[1537] Data processing: Parsing responses and extracting URLs.

[1538] Output: Initialization of the video player and streaming playback

[1539] Step 9:

[1540] Users learn sign language by watching sign language videos played on their devices and imitating the hand movements. Specifically, they practice the sign for "hello" while watching the videos.

[1541] Input: Watching sign language videos

[1542] Data calculation: Imitation of hand movements

[1543] Output: Learning sign language

[1544] Step 10:

[1545] After watching a video, the device automatically displays information about volunteer activities to the user. For example, it might display a pop-up asking, "Would you like to participate in volunteer activities?"

[1546] Input: Video viewing end event

[1547] Data processing: Displaying volunteer guidance messages

[1548] Output: Display of a popup message

[1549] Step 11:

[1550] Users follow the instructions and register to participate in volunteer activities that interest them. Specifically, they participate in an activity by filling in the required information on a form and clicking the registration button.

[1551] Input: User-generated registration information

[1552] Data processing: Form data submission

[1553] Output: Registration for volunteer activities

[1554] Through the steps outlined above, this system efficiently supports sign language learning and promotes users' participation in volunteer activities.

[1555] (Application Example 1)

[1556] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1557] In modern factory and other work environments, communication between hearing-impaired workers and machines or robots remains a challenge. Furthermore, in environments where understanding sign language is insufficient, smooth communication using sign language is difficult, potentially reducing the work efficiency of hearing-impaired workers. Additionally, a lack of motivation among the general public to learn sign language and participate in volunteer activities is also a problem.

[1558] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1559] In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in a database, means for the server to invoke a generative artificial intelligence to generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to display volunteer activity information, means for capturing the user's sign language movements using a camera and analyzing them, means for sending the analyzed sign language movements to the server and recognizing the instruction content, and means for the robot to perform an action based on the recognized instruction content. This enables smooth communication between workers with hearing impairments and robots in the factory, and makes it possible to promote sign language learning and motivate participation in volunteer activities.

[1560] A "search box" is an input field used by users to enter specific words or phrases via their device.

[1561] A "search request" is a request generated by the device based on data entered by the user and sent to the server.

[1562] A "server" is a computer system that receives and processes search requests.

[1563] A "database" is an information system for systematically storing and managing data such as sign language words and their corresponding videos, in a searchable format.

[1564] "Generative artificial intelligence" refers to artificial intelligence algorithms that can generate new sign language videos based on existing data.

[1565] A "sign language video" is a video recording of actions that express specific sign language words or phrases.

[1566] "Streaming playback" is a method of playing video data in real time, allowing users to watch before the data is fully downloaded.

[1567] A "volunteer activity guide" is a notification or message that provides users with information about volunteer activities and encourages them to participate.

[1568] A "camera" is an imaging device used to capture the user's sign language movements.

[1569] "Analysis" is the process of deciphering sign language movements captured by a camera and determining whether they represent specific sign language words or phrases.

[1570] A "robot" is a mechanical device that autonomously performs specific tasks in factories and other settings, operating based on user instructions.

[1571] A "user" is a person who uses this system to learn sign language or to give instructions to the robot through sign language.

[1572] This invention provides a system that makes it easier for users to learn sign language and enables smooth communication with robots using sign language within a factory. This system is primarily implemented using a user terminal, server, database, camera, and robot.

[1573] User terminal functions

[1574] User devices include smartphones, tablets, and personal computers. These devices are equipped with a search box where users enter sign language words or phrases they wish to learn. A search request is generated and sent to the server. The device has the capability to convert the request into JSON format.

[1575] Server Functions

[1576] The server receives a search request from the user and extracts the words from the request. Next, it searches the database for videos of the corresponding sign language words. If no video exists in the database, it calls upon a generative artificial intelligence to generate a sign language video. The generated video is also stored in the database as a cache for future searches. The server returns the search results to the user's terminal and provides the URL of the sign language video.

[1577] Camera functions

[1578] The camera is used to capture the user's sign language movements. The captured video is analyzed by an internal or external sign language recognition algorithm, and the content of the sign language is extracted. The extracted content is then sent back to the server to be used as instructions for the robot's actions.

[1579] Robot functions

[1580] The robots in the factory perform actions based on sign language instructions received from a server. For example, if a user signs "start work," the robot recognizes the sign and begins the specified task.

[1581] Specific Examples and Generative AI Models

[1582] As a concrete example of this system's use, consider a scenario where a user gives the sign language instruction to "stop work." A camera captures the sign language movement, analyzes the data, and sends it to a server. The server analyzes the data and sends the corresponding instruction to the robot. As a result, the robot immediately stops working. Additionally, when a user searches for sign language videos to learn, the generative artificial intelligence generates new sign language videos and saves them to the database.

[1583] Examples of prompt sentences to use for sign language recognition in a generative AI model are as follows:

[1584] "Design and implement a method to recognize specific sign language actions (e.g., 'start work') from short video clips."

[1585] This system will allow hearing-impaired employees to learn sign language and give natural sign language instructions to robots within the factory, which is expected to improve the working environment and increase operational efficiency.

[1586] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1587] Step 1:

[1588] The user enters the sign language word they want to learn into the search box via their device. The entered word is sent to the device as request data. Specifically, the user enters "hello" and presses the search button. In this case, the input is "hello," and it is recognized by the device.

[1589] Step 2:

[1590] The device generates a search request, converts it to JSON format, and sends it to the server. The generated JSON data is in the format {"query": "こんにちは"} and is sent to the server using the HTTP POST method. The input in this case is the word "こんにちは" entered by the user, and the output is data in JSON format.

[1591] Step 3:

[1592] The server receives a search request and extracts a word from the request body. Specifically, it extracts the value "こんにちは" from the "query" field of the received JSON data. The input in this case is JSON data {"query": "こんにちは"}, and the output is the extracted word "こんにちは".

[1593] Step 4:

[1594] The server searches the database for the corresponding sign language word. Specifically, it uses a database query to search for a sign language video corresponding to "hello." The input in this case is the extracted word "hello," and the output is the URL of the corresponding sign language video.

[1595] Step 5:

[1596] If the server does not find a corresponding sign language video in the database, it invokes a generative artificial intelligence to generate one. The generated sign language video is then saved as a cache in the database. The input in this case is the extracted word "hello," and the output is the generated sign language video.

[1597] Step 6:

[1598] The server returns the search results to the user's device. The search results include URLs for sign language videos. Specifically, the URLs are constructed as a JSON response and sent to the user's device. The input in this process is the search results or the URLs of the generated sign language videos, and the output is the JSON response sent to the user's device.

[1599] Step 7:

[1600] The terminal processes the response received from the server, extracts the URL of the sign language video, and starts streaming playback. Specifically, it initializes the video player and plays the sign language video. The input in this process is a JSON response, and the output is the streaming playback of the sign language video.

[1601] Step 8:

[1602] Users learn sign language by watching sign language videos on their devices. Specifically, they imitate hand movements while watching the videos being played. In this process, the input is the sign language video, and the output is the learning of sign language.

[1603] Step 9:

[1604] The device uses its camera to capture the user's sign language movements and analyzes them. Specifically, it analyzes the video data captured by the camera using an internal algorithm to extract the content of the sign language. The input in this process is the captured sign language video, and the output is the extracted sign language content.

[1605] Step 10:

[1606] The terminal analyzes the sign language gestures and sends them to the server, which then recognizes the instructions. The analysis results are converted to JSON format and sent using the HTTP POST method. The input in this process is the extracted sign language content, and the output is the transmission of JSON data to the server.

[1607] Step 11:

[1608] The robot executes actions based on the instructions recognized by the server. Specifically, the robot receives recognition data from the server and performs the corresponding action, such as "start work." In this process, the input is the recognized instruction, and the output is the robot's action.

[1609] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1610] This invention relates to a system for promoting sign language learning and participation in volunteer activities, with the aim of supporting people with hearing impairments. Furthermore, this invention aims to provide users with an optimal learning experience by combining it with an emotion engine that recognizes the user's emotions. The processing flow of the system's program is described below in natural language and illustrated in detail with specific examples.

[1611] System processing flow

[1612] 1. User search for sign language words:

[1613] The user enters the sign language word they want to learn into the search box via their device and clicks the search button. For example, they might enter "hello".

[1614] 2. Sending request data via the terminal:

[1615] The terminal receives user input data, generates a search request in JSON format, and sends it to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[1616] 3. Receiving and parsing requests by the server:

[1617] The server receives the HTTP request, extracts the search term "hello" from the request body, and generates the query necessary for the database search.

[1618] 4. Database search by server:

[1619] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[1620] 5. Server-based invocation of generative artificial intelligence:

[1621] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video. This generated video is stored in the database as a cache for future searches.

[1622] 6. Server sends response data:

[1623] The server constructs response data containing the URL of the sign language video and sends it back to the terminal. For example, it sends data like {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"} to the terminal.

[1624] 7. Processing and display of response data by the terminal:

[1625] The terminal processes the response received from the server, extracts the URL of the sign language video, initializes the video player, and streams the sign language video.

[1626] 8. User learning of sign language:

[1627] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements. For example, they can practice the sign for "hello" while watching a video.

[1628] 9. User emotion recognition by emotion engine:

[1629] The device captures the user's facial expressions through its camera, and an emotion engine analyzes those expressions to recognize the user's emotions. For example, if the user is smiling during the learning process, it will be recognized as a positive emotion.

[1630] 10. Adjusting the learning content:

[1631] Based on the user's emotions recognized by the emotion engine, the device adjusts the learning content. For example, if it detects a confused expression on the user's face, it will provide supplementary learning materials or explanations.

[1632] 11. Recommendation of appropriate sign language learning content:

[1633] The device recommends appropriate sign language learning content based on the user's emotional state. For example, if the user is in a positive emotional state, it will suggest more difficult sign language.

[1634] 12. Display of volunteer activity information on terminals:

[1635] After learning sign language, the device automatically displays information about volunteer activities to the user. For example, a pop-up message such as "Would you like to participate in volunteer activities?" might appear.

[1636] 13. User registration for volunteer activities:

[1637] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[1638] As described above, the system of the present invention efficiently supports sign language learning and provides a more effective learning experience by adjusting the learning content according to the user's emotions. Furthermore, it can enhance social contribution by promoting participation in volunteer activities.

[1639] The following describes the processing flow.

[1640] Step 1:

[1641] The user enters the sign language word they want to learn into the search box on their device and clicks the search button. For example, they might enter "hello".

[1642] Step 2:

[1643] The terminal receives user input data and generates a search request in JSON format. The generated JSON request is sent to the server using the HTTP POST method. For example, {"query": "Hello"} is sent to the server.

[1644] Step 3:

[1645] The server receives an HTTP request. After receiving it, it extracts the search term "hello" from the request body and generates the query necessary for the database search. Specifically, it parses the request and identifies the word entered by the user.

[1646] Step 4:

[1647] The server generates a query to search the database for the corresponding sign language word "hello". For example, the SQL query SELECT FROM sign_language WHERE word='こんにちは'; is executed.

[1648] Step 5:

[1649] The server checks the search results from the database. If a matching sign language video exists in the database, it retrieves the URL of that video. If it does not exist, it proceeds to the next step.

[1650] Step 6:

[1651] The server invokes a generative artificial intelligence to generate a sign language video corresponding to the sign language word "hello" requested by the user. This generated video is stored as a cache in the database for future searches.

[1652] Step 7:

[1653] The server constructs response data that includes the URL of the sign language video. For example, it generates data in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[1654] Step 8:

[1655] The server sends the generated response data back to the user's terminal. The terminal receives the HTTP response.

[1656] Step 9:

[1657] The terminal analyzes the response received from the server and extracts the URL of the sign language video. Then, it initializes the video player and prepares it for streaming playback of the sign language video.

[1658] Step 10:

[1659] Users watch sign language videos on their devices and learn sign language by imitating the hand movements. Specifically, they practice the sign language for "hello" while watching the videos.

[1660] Step 11:

[1661] The device uses its camera to capture the user's facial expressions, and an emotion engine analyzes those expressions to recognize the user's emotions. For example, if the user is confused, the system analyzes and understands that expression.

[1662] Step 12:

[1663] Based on the user's emotions recognized by the emotion engine, the device adjusts the learning content. Specifically, if the user is confused, it displays supplementary materials or additional explanations.

[1664] Step 13:

[1665] The device recommends appropriate sign language learning content based on the user's emotional state. For example, if the user is in a positive state, it will suggest more difficult sign language.

[1666] Step 14:

[1667] After watching a sign language video, the device automatically displays information about social activities and volunteer opportunities to the user. For example, a pop-up message might appear asking, "Would you like to participate in volunteer activities?"

[1668] Step 15:

[1669] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[1670] (Example 2)

[1671] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1672] Traditional sign language learning systems had problems such as the inability for users to easily obtain videos corresponding to the sign language words they searched for, which reduced the efficiency of sign language learning. Furthermore, the lack of customization of learning content based on the user's emotional state meant that learning effectiveness was not maximized, potentially leading to decreased user motivation. In addition, there was a lack of means to encourage participation in volunteer activities, resulting in fewer opportunities to raise awareness of social contribution.

[1673] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in the database, means for the server to call a generative artificial intelligence and generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to capture the user's facial expression via a camera and for the emotion engine to recognize the user's emotion, means for the terminal to adjust the learning content based on the recognized emotion, and means for the terminal to display volunteer activity information. As a result, the user can efficiently proceed with learning sign language, adjust the learning content based on their emotional state, and further increase their motivation to participate in volunteer activities.

[1674] A "terminal" is a device that a user directly operates to input and display information.

[1675] A "server" is a device that processes data sent from a terminal and provides the necessary information.

[1676] A "search box" is a graphical user interface element that users use to enter specific information.

[1677] A "search request" is a request made to query a server for necessary data based on information entered by the user.

[1678] "JSON format" is an abbreviation for JavaScript Object Notation, and it is a lightweight text data exchange format for structuring and representing data.

[1679] The "HTTP POST method" is an HTTP request method used by a client to send resources to a server.

[1680] An "emotion engine" is software or a device used to recognize and analyze a user's emotions.

[1681] A "sign language video" is a video recording of sign language movements and expressions.

[1682] "Generative artificial intelligence" is an artificial intelligence technology that has the ability to generate new content based on a given prompt.

[1683] A "prompt message" is text data input to a generative artificial intelligence system to generate a specific response.

[1684] "Streaming playback" is a technology that allows playback while receiving data in real time.

[1685] "Volunteer Activity Guide" refers to guide information displayed to users to obtain information about volunteer activities and participate in them.

[1686] Modes for carrying out the invention

[1687] This invention relates to a sign language learning system for supporting people with hearing impairments and a system for promoting participation in volunteer activities. The system aims to provide an optimal learning experience by recognizing the user's emotions. The following describes in detail the embodiments for specifically carrying out this invention.

[1688] The user enters the sign language word they want to learn into the search box via their device and clicks the search button. Upon receiving this input, the device generates a search request in JSON format and sends it to the server using the HTTP POST method. For example, if the user wants to learn the sign language word for "hello," they would search for "hello" in the input, and the device would generate JSON data {"query": "hello"} and send it to the server.

[1689] The server receives this HTTP request and extracts the search word from the request body. Based on the extracted word, the server generates an SQL query to search for the corresponding sign language word in the database and executes the database search. For example, it executes the SQL query SELECT FROM sign_language WHERE word='こんにちは';

[1690] If the database search does not find a matching sign language word, the server calls a generative artificial intelligence (generative AI model) and generates a sign language video based on the prompt text. This generated sign language video is stored in the database as a cache for future searches. For example, the generative AI model will generate a sign language video if the prompt text "Please generate a sign language video for 'hello'" is entered into it.

[1691] A response data containing the URL of the generated or retrieved sign language video is constructed, and the server returns it to the terminal in JSON format. For example, it might be returned in the format {"word": "こんにちは", "video_url": "http: / / example.com / signs / こんにちは.mp4"}.

[1692] The device processes the received response data, extracts the URL of the sign language video, and initializes the video player. It then streams the sign language video, allowing the user to watch and learn from it on their device.

[1693] Furthermore, the device captures the user's facial expressions via the camera, and an emotion engine analyzes this facial data to recognize the user's emotions. For example, if the user is learning sign language with a smile, a positive emotion will be recognized. Based on this emotion data, the device adjusts the learning content. For instance, if the user shows a confused expression, the device will provide additional learning materials or clearer explanations.

[1694] After completing sign language learning, the device automatically displays information about volunteer activities and encourages the user to participate. For example, it might display a pop-up asking, "Would you like to participate in volunteer activities?" The user can then register to participate by following this prompt.

[1695] This system allows users to learn sign language efficiently, adjust learning content according to their emotions, and foster a sense of social responsibility by encouraging participation in volunteer activities.

[1696] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1697] Step 1:

[1698] The user enters the sign language word they want to learn into the search box via their device and clicks the search button.

[1699] Input: Sign language words entered by the user into the search box on the device (e.g., "hello")

[1700] Output: Clicking the search button sends the entered sign language word as a search request to the terminal.

[1701] Step 2:

[1702] The terminal receives user input data, generates a search request in JSON format, and sends it to the server using the HTTP POST method.

[1703] Input: Sign language word (e.g., "hello")

[1704] Data processing: Generate JSON data containing sign language words (e.g., {"query": "Hello"}).

[1705] Output: JSON data is sent to the server using the HTTP POST method.

[1706] Step 3:

[1707] The server receives an HTTP request and extracts the search term from the request body.

[1708] Input: JSON data sent to the server

[1709] Data processing: Extraction of search terms (e.g., "hello")

[1710] Output: Extracts search terms and prepares a query for database search.

[1711] Step 4:

[1712] The server searches the database for the corresponding sign language word.

[1713] Input: Search term (e.g., "hello")

[1714] Data calculation: Executing database queries (e.g., SELECT FROM sign_language WHERE word='こんにちは';)

[1715] Output: Search results (If a relevant sign language video is found, that will be the result; otherwise, proceed to the next step.)

[1716] Step 5:

[1717] If a corresponding sign language word is not found in the database, the server uses generative artificial intelligence to generate a sign language video.

[1718] Input: Search term if no matching sign language word was found.

[1719] Data processing: Input a prompt message into a generative artificial intelligence (e.g., "Please generate a sign language video of 'hello'").

[1720] Output: Generated sign language video

[1721] Step 6:

[1722] The server saves the generated sign language videos as a cache in the database.

[1723] Input: Generated sign language video

[1724] Data processing: Caching of sign language videos and their metadata

[1725] Output: Sign language videos cached in the database

[1726] Step 7:

[1727] The server constructs response data containing the URL of the sign language video and sends it to the terminal in JSON format.

[1728] Input: URL of the sign language video (Example: "http: / / example.com / signs / こんにちは.mp4")

[1729] Data processing: Construction of response data (Example: {"word": "Hello", "video_url": "http: / / example.com / signs / Hello.mp4"})

[1730] Output: Response data is sent to the terminal.

[1731] Step 8:

[1732] The terminal processes the response received from the server, extracts the URL of the sign language video, and initializes the video player.

[1733] Input: Response data (Example: {"word": "Hello", "video_url": "http: / / example.com / signs / Hello.mp4"})

[1734] Data processing: URL extraction and video player initialization.

[1735] Output: Start streaming the sign language video in the video player.

[1736] Step 9:

[1737] Users watch sign language videos played on their devices and learn sign language by imitating the hand movements.

[1738] Input: Sign language video (Example: Sign language video for "hello")

[1739] Action: Practice sign language while watching a video.

[1740] Output: User's proficiency in sign language improves.

[1741] Step 10:

[1742] The device captures the user's facial expressions via its camera, and an emotion engine analyzes those expressions to recognize the user's emotions.

[1743] Input: User's facial expression

[1744] Data processing: Analysis of facial expression data using an emotion engine.

[1745] Output: User sentiment data

[1746] Step 11:

[1747] Based on the user's emotions recognized by the emotion engine, the device adjusts its learning content.

[1748] Input: User sentiment data

[1749] Data processing: Adjusting learning content based on emotional data (e.g., providing supplementary materials)

[1750] Output: Adjusted learning content is provided.

[1751] Step 12:

[1752] The device recommends appropriate sign language learning content based on the user's emotional state.

[1753] Input: User sentiment data

[1754] Data processing: Evaluation of training data and determination of recommendations.

[1755] Output: Recommended sign language learning content

[1756] Step 13:

[1757] After learning sign language, the device automatically displays information about volunteer activities to the user.

[1758] Input: Learning completed

[1759] Data processing: Generating volunteer activity guides

[1760] Output: Information about volunteer activities will be displayed.

[1761] Step 14:

[1762] The user follows the instructions for the volunteer activity and proceeds to the registration page. The user fills in the required information on the form and clicks the registration button to complete their participation.

[1763] Input: Volunteer activity information

[1764] Operation: User registration

[1765] Output: Participation in volunteer activity is complete.

[1766] (Application Example 2)

[1767] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1768] In modern brick-and-mortar stores, there is a lack of support for hearing-impaired individuals to shop smoothly, particularly in the provision of product information and guidance in sign language. Furthermore, there are insufficient means of providing emotionally appropriate learning content to sign language learners. As a result, hearing-impaired individuals and sign language learners are often dissatisfied with their shopping experiences in physical stores.

[1769] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for the user to input into a search box via a terminal, means for the terminal to generate a search request and send it to the server, means for the server to receive the search request and extract the entered word, means for the server to search for the corresponding sign language word in the database, means for the server to call a generative artificial intelligence and generate a sign language video, means for the server to return the search results to the user's terminal, means for the terminal to process the response received from the server and stream the sign language video, means for the user to watch the sign language video on the terminal and learn sign language, means for the terminal to recognize the user's emotions and adjust the learning content, means for the terminal to provide guidance and product information in sign language within a physical store, and means for the terminal to display information about volunteer activities. This makes it possible for people with hearing impairments to smoothly obtain information and shop in physical stores, and for sign language learners to effectively advance their learning.

[1770] "Terminal" refers to a device operated by a user, and includes smartphones, smart glasses, head-mounted displays, or robots.

[1771] A "search box" is an input field where users enter specific words or phrases.

[1772] A "server" is a computer system that receives search requests and performs processing such as accessing databases and using generative artificial intelligence.

[1773] "Sign language words" are specific hand shapes and movements used by people with hearing impairments to communicate.

[1774] "Generative artificial intelligence" is an artificial intelligence technology that generates new sign language videos based on specific input data.

[1775] A "sign language video" is a video that visually displays sign language words and is used to help people with hearing impairments understand them.

[1776] "Emotion recognition" is a technology that captures a user's facial expressions and movements through cameras and sensors and analyzes their emotional state.

[1777] "Adjusting learning content" refers to dynamically changing the difficulty level and content of the learning materials and sign language videos provided according to the user's emotional state.

[1778] A "physical store" is a commercial facility that exists physically, a place where customers can visit in person and purchase goods.

[1779] "Volunteer activities" refer to support activities using sign language and other social contribution activities, and are acts in which users contribute to society through their participation.

[1780] "Response" refers to response data such as search results and sign language videos sent from the server to the terminal.

[1781] This invention relates to a system designed to support the hearing impaired and sign language learners, and is particularly intended to provide users with sign language guidance and product information in physical stores. The system also has a function to recognize the user's emotions and adjust the learning content accordingly.

[1782] Hardware and software to be used

[1783] The hardware for carrying out the present invention includes smart glasses, a camera, and a microphone. As a specific example, the smart glasses are suitable devices such as Google Glass. The software includes an emotion recognition engine (e.g., Microsoft Azure Face API), a sign language generation AI model, a cloud database (e.g., Google Firebase), and a real-time translation API (e.g., Azure Translator).

[1784] System Configuration

[1785] 1. Terminal Operation: The user wears smart glasses and enters sign language words or instructions into the search box. Input can be done via voice recognition, touchpad, or gestures.

[1786] 2. Sending a search request: The device generates a search request, converts it to JSON format, and sends it to the server. For example, the query "Where is the water?" is sent.

[1787] 3. Server Processing: The server receives the request and extracts the entered word. It then searches the database for the corresponding sign language word. If no search result is found, it invokes a generative artificial intelligence to generate a new sign language video and saves it as a cache.

[1788] 4. Sending the response: The server constructs response data that includes the search results and sends it back to the terminal. For example, it may include a URL for a sign language video of "water".

[1789] 5. Processing the response and playing the video: The terminal processes the received response, extracts the URL of the sign language video, and streams it.

[1790] 6. Emotion Recognition and Learning Content Adjustment: The device captures the user's facial expressions through the camera and analyzes them with an emotion recognition engine. For example, if the system recognizes that the user is confused, it will provide supplementary explanations or additional learning materials.

[1791] 7. In-store guidance: Users can view sign language videos through smart glasses while receiving guidance and product information within a physical store. This allows people with hearing impairments to shop more smoothly.

[1792] 8. Volunteer Activity Information: Finally, the device will display information about volunteer activities after the sign language learning session. If the user is interested, detailed information will be provided, and they can easily register to participate.

[1793] Specific example

[1794] Examples of prompts to input into a generative AI model:

[1795] Please create a video of the sign language word for "water." Clearly demonstrate the hand shape and movement used.

[1796] Thus, this system provides effective support for the hearing impaired and sign language learners, significantly improving the in-store experience. By dynamically adjusting learning content according to the user's emotions, it can provide a more personalized learning experience.

[1797] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1798] Step 1:

[1799] The user wears smart glasses and enters a search query into the search box. Input methods include voice recognition, touchpad, or gesture input. This query may refer to product information or guidance the user wants to know. For example, the user might enter the query, "Where is the water?"

[1800] Input: Query (Example: "Where is the water?")

[1801] Output: Text data of the search query

[1802] Step 2:

[1803] The terminal receives the search query and converts it into a JSON-formatted search request. For example, the query "Where is the water?" is converted to the format {"query": "Where is the water?"}.

[1804] Input: Query text data

[1805] Output: Search request in JSON format

[1806] Step 3:

[1807] The terminal sends the generated search request in JSON format to the server. The data is sent to the server using the HTTP POST method.

[1808] Input: Search request in JSON format

[1809] Output: Request sent to the server

[1810] Step 4:

[1811] The server receives the HTTP request and extracts the search query (word) from the request body. In this example, the word "water" is extracted.

[1812] Input: Request sent to the server

[1813] Output: Extracted word (e.g., "water")[[]END]]

[1814] Step 5:

[1815] The server uses the extracted word to search for the corresponding sign language word in the database. The SQL query SELECT FROM sign_language WHERE word='water'; is executed to retrieve the data corresponding to the sign language word "water" from the database.

[1816] Input: Extracted word

[1817] Output: Search result (data of sign language word)

[1818] Step 6:

[1819] If the server does not find data corresponding to the sign language word in the database, it calls a generative artificial intelligence to generate a new sign language video. The generated sign language video is saved as a cache in the database for future searches.

[1820] Input: Data of sign language word (if not present)

[1821] Output: Generated sign language video data

[1822] Step 7:

[1823] The server constructs response data including the search result or the URL of the generated sign language video and returns it to the terminal. For example, data like {"word": "water", "video_url": "http: / / example.com / signs / water.mp4"} is returned.

[1824] Input: Data of search result or generated sign language video

[1825] Output: Response data

[1826] Step 8:

[1827] The terminal processes the response data received from the server and extracts the URL of the sign language video. Then, it initializes the video player on the terminal and streams and plays the sign language video.

[1828] Input: Response data

[1829] Output: Streaming playback of sign language video

[1830] Step 9:

[1831] The user watches the sign language video played on the terminal and imitates the hand movements to learn sign language. The user practices the sign language for "water" while watching the video.

[1832] Input: Streaming playback of sign language videos

[1833] Output: Learning sign language

[1834] Step 10:

[1835] The device captures the user's facial expressions through its camera and analyzes the user's emotions using an emotion recognition engine (Microsoft Azure Face API). For example, if it captures a confused expression from the user, it analyzes this information.

[1836] Input: Captured user facial expression

[1837] Output: User's emotional state

[1838] Step 11:

[1839] The device adjusts the learning content based on the user's emotional state. If the user is confused, it provides supplementary explanations and additional materials; if the user is in a positive emotional state, it suggests more difficult sign language.

[1840] Input: User's emotional state

[1841] Output: Adjusted learning content

[1842] Step 12:

[1843] After the user finishes learning sign language, the device will display information about volunteer activities. For example, a message such as "Would you like to participate in volunteer activities?" will be displayed.

[1844] Input: Sign language learning completed

[1845] Output: Display of volunteer activity information

[1846] The above outlines the specific processing steps of how the system of this invention actually works. This system will enable people with hearing impairments and sign language learners to have a better experience in physical stores.

[1847] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1848] The data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1849] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1850] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1851] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1852] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1853] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1854] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1855] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1856] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1857] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1858] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1859] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1860] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1861] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1862] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1863] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1864] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1865] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1866] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1867] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[1868] The following is further disclosed regarding the embodiments described above.

[1869] (Claim 1)

[1870] The means by which the user enters information into the search box via the device,

[1871] A means by which the terminal generates a search request and sends it to the server,

[1872] The server receives a search request and extracts the entered words,

[1873] The server has a means of searching for the corresponding sign language word in the database,

[1874] A means by which a server calls a generative artificial intelligence to generate sign language videos,

[1875] A means by which the server returns search results to the user's terminal,

[1876] A means for the terminal to process the response received from the server and stream a sign language video,

[1877] A means for users to learn sign language by watching sign language videos on their devices,

[1878] A means for the terminal to display volunteer activity information,

[1879] A system that includes this.

[1880] (Claim 2)

[1881] When the device generates a search request, it includes a means of converting the request to JSON format.

[1882] The system according to claim 1.

[1883] (Claim 3)

[1884] The server includes means for storing search results as a cache in the database.

[1885] The system according to claim 1.

[1886] "Example 1"

[1887] (Claim 1)

[1888] The means by which the user enters information into the search box via the device,

[1889] A means by which the terminal generates a search request and sends it to the server,

[1890] A server receives a search request and extracts the entered search terms,

[1891] The server has a means of searching for the corresponding sign language word in the database,

[1892] If the server does not find a corresponding sign language word in the database, it invokes a generative artificial intelligence to generate a sign language video.

[1893] A method for saving sign language videos generated using generative artificial intelligence as a cache in a database,

[1894] A means by which the server returns search results to the user's terminal,

[1895] A means for the terminal to process the response received from the server and stream a sign language video,

[1896] A means for users to learn sign language by watching sign language videos on their devices,

[1897] A means for the terminal to display volunteer activity information,

[1898] A system that includes this.

[1899] (Claim 2)

[1900] When the device generates a ...

Claims

1. The means by which the user enters information into the search box via the device, A means by which the terminal generates a search request and sends it to the server, The server receives a search request and extracts the entered words, The server has a means of searching for the corresponding sign language word in the database, A means by which a server calls a generative artificial intelligence to generate sign language videos, A means by which the server returns search results to the user's terminal, A means for the terminal to process the response received from the server and stream a sign language video, A means for users to learn sign language by watching sign language videos on their devices, A means for the terminal to display volunteer activity information, A system that includes this.

2. When the device generates a search request, it includes a means of converting the request to JSON format. The system according to claim 1.

3. The server includes means for storing search results as a cache in the database. The system according to claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A