system
The system enhances sign language learning by analyzing user input through a server to provide background information and meaning, addressing regional diversity and improving educational impact.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-03-16
AI Technical Summary
Conventional sign language learning methods face challenges due to regional and age-related diversity, leading to a lack of unified understanding, reduced interest in learning, and insufficient educational impact.
A system that allows users to input sign language data, which is analyzed by a server to retrieve background information and meaning using image recognition algorithms, and display the results on a terminal, enhancing understanding and familiarity.
Improves the efficiency and accessibility of sign language learning by providing detailed background information and meaning, making it easier for users to understand and apply sign language effectively.
Smart Images

Figure 2026047938000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional sign language learning methods, there is diversity in sign languages that vary depending on region and age, making it difficult to obtain a unified understanding. For this reason, it is difficult to have a sense of familiarity with sign language, and there is a problem that the number of people who want to learn sign language decreases. In addition, there are few opportunities to deeply understand the background and meaning of sign language, and the educational effect of sign language has also declined.
Means for Solving the Problems
[0005] This invention provides a system in which a user inputs sign language and sends that sign language data to a server. The server analyzes the received sign language data and retrieves background information and meaning related to the sign language from a database. The retrieved background information and meaning are then sent from the server to a terminal, which displays them to the user. This makes it easier for the user to accurately understand the background and meaning of the sign language. Furthermore, by using an image recognition algorithm, it is possible to analyze the sign language data with high accuracy and associate background information and meaning using specific identifiers.
[0006]
[0007] A "user" is an individual or group that uses the system to input sign language and receives the results.
[0008] Sign language is a visual language based on hand and body movements, primarily used by people with hearing impairments to communicate.
[0009] "Means of input" refers to methods or devices for capturing sign language as digital data using a camera, microphone, or existing video files.
[0010] A "server" is a computer system that receives sign language data via a network, analyzes it, retrieves the results from a database, and transmits them.
[0011] "Means of transmission" refers to a method or device for sending data to a server via the Internet or other communication networks.
[0012] "Means of analysis" refer to algorithms and software used to analyze sign language data and identify its meaning and related information.
[0013] "Background information" refers to information about the cultural and historical background of how that sign language came into being.
[0014] "Meaning" refers to information about the content and intended message expressed by sign language.
[0015] A "database" is a digital information storage system that stores background information and meanings related to sign language, and allows for the retrieval of that information.
[0016] A "terminal" refers to a device such as a computer, smartphone, or tablet used by a user to input sign language and display information retrieved from a server.
[0017] An "image recognition algorithm" is a mathematical method or program used to analyze image data, such as camera footage, and extract specific patterns or features from it.
[0018] An "identifier" is a unique number or symbol used to distinguish a particular sign language from other sign languages. [Brief explanation of the drawing]
[0019] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8]It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
Mode for Carrying Out the Invention
[0020] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0021] First, the language used in the following description will be explained.
[0022] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0023] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0027] [First Embodiment]
[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0040] This invention relates to a system for learning sign language in a more accessible way, and in particular, a system that helps users input sign language and understand its background information and meaning. The system consists of three main elements: the user, the terminal, and the server.
[0041] Program Processing Overview
[0042] 1. The user enters the sign language.
[0043] Users perform sign language using their device's camera. Alternatively, they can upload pre-recorded sign language video files to their device.
[0044] 2. The device sends sign language data to the server.
[0045] The device saves the sign language video captured by its camera as digital data. Next, it transmits this sign language data to a server via the internet.
[0046] 3. The server analyzes the sign language data.
[0047] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Specifically, it extracts characteristic parts of the sign language and compares them with an existing sign language dictionary.
[0048] 4. The server retrieves background information and meaning.
[0049] The server queries the database using the identified sign language identifier to retrieve background information and meanings for that sign language. For example, it may retrieve information about the origins of the sign language, as well as its cultural and historical background.
[0050] 5. The server sends information to the terminal.
[0051] The server converts the acquired background information and its meaning into structured data and sends it to the terminal.
[0052] 6. The device displays the information.
[0053] The terminal receives background information and meaning from the server and displays it to the user. Display formats include text, images, and videos.
[0054] Specific example
[0055] Example 1: Inputting the sign language for "thank you"
[0056] The user performs the sign language for "thank you" while pointing it at the device's camera. The device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as the sign language for "thank you." The server retrieves background information and meaning of the sign language for "thank you" from its database and sends it to the device. The device then displays this information to the user. Specifically, it might display something like, "The sign language for 'thank you' is used to express gratitude in Japanese."
[0057] Example 2: Inputting regional sign language
[0058] The user inputs sign language specific to the Kansai region into the terminal. The terminal captures this sign language and sends it to the server. The server analyzes the sign language data and recognizes that this sign language is specific to the Kansai region. The server retrieves the regional background and cultural meaning of this sign language from its database and sends it to the terminal. The terminal displays this information to the user. Specifically, it displays, "This sign language represents a greeting unique to people in the Kansai region."
[0059] In this way, by providing users with background information and meanings of sign language data, this system can improve the efficiency of sign language learning and enhance users' familiarity with sign language.
[0060] The following describes the processing flow.
[0061] Step 1: The user enters sign language.
[0062] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[0063] Step 2: The device captures and saves the sign language data.
[0064] The device saves sign language videos captured by the camera as digital data. Similarly, uploaded video files are also saved as digital data.
[0065] Step 3: The device sends the sign language data to the server.
[0066] The terminal transmits the stored sign language data to the server via the internet. During this process, appropriate pre-processing, such as formatting and compression, is performed on the data.
[0067] Step 4: The server receives the sign language data.
[0068] The server receives sign language data sent from the terminal.
[0069] Step 5: The server analyzes the sign language data.
[0070] The server applies an image recognition algorithm to analyze the received sign language data. It extracts characteristic parts of sign language movements and evaluates their similarity by comparing them with an existing sign language dictionary.
[0071] Step 6: The server identifies the sign language identifier.
[0072] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database.
[0073] Step 7: The server retrieves background information and meaning.
[0074] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[0075] Step 8: The server converts background information and meaning into structured data.
[0076] The server converts the acquired background information and meaning into structured data in a format that is easy for the user to read.
[0077] Step 9: The server sends the structured data to the terminal.
[0078] The server sends structured data containing background information and semantics to the terminal.
[0079] Step 10: The device receives and displays information.
[0080] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, videos, etc., and is provided in a way that makes it easy for the user to understand the background information and meaning of sign language.
[0081] In this way, users can understand detailed background information and meaning related to the sign language they input. This process is designed to make learning sign language efficient and accessible.
[0082] (Example 1)
[0083] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0084] When learning sign language, many users face the challenge of understanding the precise meaning and background information of sign language movements. Furthermore, existing sign language learning systems often suffer from low accuracy in analyzing sign language or insufficient acquisition of background information, which hinders the improvement of sign language learning efficiency.
[0085] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0086] In this invention, the server includes means for analyzing transmitted sign language data using a machine learning algorithm, means for obtaining background information and meaning related to the analyzed sign language data from a data store, and means for sending the obtained background information and meaning to a terminal. As a result, when a user inputs sign language, they can quickly obtain the accurate meaning and background information of that sign language.
[0087] A "user" is a person who uses this system to input sign language and receives the analysis results and background information.
[0088] An "input device" is a device used by a user to input sign language, and specifically refers to a camera or touchscreen.
[0089] A "communication device" refers to a device used to transmit input sign language data to a server, and specifically includes wireless communication modules and internet connectivity functions.
[0090] A "server" is a computing device responsible for analyzing sign language data, acquiring background information and meaning, and transmitting this information to users.
[0091] A "machine learning algorithm" is a mathematical method used to analyze sign language data and identify its meaning, specifically referring to Convolutional Neural Networks (CNNs), among others.
[0092] A "data store" refers to a database or storage system used to store background information and meanings of sign language.
[0093] A "display device" is a device used to visually display received information to a user, and specifically refers to displays and monitors.
[0094] A "unique identifier" is a code or number used to uniquely identify a particular sign language, and is used in database queries.
[0095] Modes for carrying out the invention
[0096] This invention relates to a system that helps users input sign language and understand its background information and meaning. The following describes specific embodiments for carrying out the invention.
[0097] Hardware and software to be used
[0098] 1. User
[0099] The user's role is to input sign language. This is done using input devices equipped with cameras and touchscreens.
[0100] 2. Terminal
[0101] The terminal's role is to capture sign language data entered by the user and send it to the server. The terminal's communication equipment includes a wireless communication module and internet connectivity.
[0102] The terminal includes video capture devices (e.g., cameras) and data display devices (e.g., displays).
[0103] 3. Server
[0104] The server's role is to analyze sign language data, extract background information and meaning, and provide it to the user. Sign language analysis utilizes image recognition algorithms such as Convolutional Neural Networks (CNNs). Specifically, machine learning frameworks like TENSORFLOW® and PyTorch are used.
[0105] The server also has a data transmission function to send the acquired background information and meaning to the terminal.
[0106] 4. Datastore
[0107] Data stores are used to store background information and meanings of sign language. SQL and NoSQL databases are typical examples.
[0108] Specific example
[0109] Example 1: Inputting "thank you" in sign language
[0110] 1. The user performs the sign language for "thank you" while pointing it at the device's camera.
[0111] The user performs the sign language for "thank you" in front of the camera.
[0112] 2. The device captures this video and sends it to the server as sign language data.
[0113] The device's camera captures the video frame by frame and saves it as digital data. The data is then sent to a server.
[0114] 3. The server analyzes the sign language data and identifies it as the sign for "thank you."
[0115] The server analyzes the received data using an image recognition algorithm and identifies it as the sign language for "thank you."
[0116] 4. The server retrieves background information and meaning of the sign language for "thank you" from the database and displays the retrieved information on the terminal.
[0117] The server retrieves information related to the sign language expression for "thank you" from its database. For example, this information may include details about how gratitude is expressed in Japanese culture. The retrieved information is then sent to the terminal as structured data (e.g., in JSON format).
[0118] 5. Display the information received by the device to the user.
[0119] The device visually displays the information it receives to the user. For example, it might display information such as, "The sign language for 'thank you' is used to express gratitude in Japanese."
[0120] Example of a prompt
[0121] The following is an example of a prompt statement for inputting a system description into the generating AI model.
[0122] Prompt message:
[0123] I would like to describe a system for learning sign language. This system helps users input sign language and understand its meaning and background information. Please explain the program's process in natural language, following the steps below.
[0124] procedure:
[0125] 1. The user inputs the sign language. The user either performs the sign language using their device's camera or uploads a pre-recorded video file of the sign language.
[0126] 2. The device sends sign language data to the server. The device saves the sign language video captured by the camera as digital data and sends it to the server via the internet.
[0127] 3. The server analyzes the sign language data. The server uses an image recognition algorithm to analyze what the sign language data means and compares it to an existing sign language dictionary.
[0128] 4. The server retrieves background information and meaning. The server queries the database using the identified sign language identifier to retrieve background information and meaning for that sign language.
[0129] 5. The server sends the information to the terminal. The server converts the acquired background information and meaning into structured data and sends it to the terminal.
[0130] 6. The terminal displays information. The terminal receives background information and meaning sent from the server and displays it to the user. Display formats include text, images, videos, etc.
[0131] Specific example:
[0132] Please explain the processing steps using the sign language for "thank you" as an example.
[0133] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0134] Specific flow of program processing
[0135] Processing steps
[0136] Step 1:
[0137] The user inputs sign language. The user either performs sign language using the camera on their device or uploads a pre-recorded video file of sign language. The input data is a video of sign language. This video of sign language becomes the raw data used for analysis in subsequent processing.
[0138] Step 2:
[0139] The device captures sign language data and saves it as digital data. The device's camera captures video frame by frame and saves this as digital data (e.g., MP4 format) to internal storage. The captured data is used for subsequent analysis.
[0140] Step 3:
[0141] The terminal transmits the stored sign language data to the server via a communication device. The terminal uses an internet connection to transmit the stored sign language data to the server. The input data is the captured sign language video data, and the output is the data transmitted to the server.
[0142] Step 4:
[0143] The server receives sign language data and analyzes it using a machine learning algorithm. The server then analyzes the received data using an image recognition algorithm (e.g., a Convolutional Neural Network). This analysis extracts the features of the sign language and converts them into an identifiable format. The input data is video data of sign language, and the output is the analyzed sign language data.
[0144] Step 5:
[0145] The server identifies a unique identifier based on the analyzed sign language data and retrieves background information and meaning from the data store. The server queries the database based on the analysis results to retrieve background information and meaning for the corresponding sign language. The input data is the analyzed sign language data, and the output is data related to background information and meaning.
[0146] Step 6:
[0147] The server converts the acquired information into structured data (e.g., JSON format) and sends it to the terminal. The server uses a data transmission function to send the acquired information to the terminal. The input data is background information and semantic data, and the output is the structured data sent to the terminal.
[0148] Step 7:
[0149] The terminal displays the information it receives to the user. The terminal's display device analyzes the received structured data and displays it visually to the user. Display formats include text, images, and videos. The input data is structured data, and the output is the information displayed to the user.
[0150] Detailed explanation of operation
[0151] Step 1:
[0152] Users perform the sign language for "thank you" using their camera. Users can also upload the recorded video file.
[0153] Step 2:
[0154] The device's camera captures video of the user performing sign language. This video is saved as digital data to the internal storage.
[0155] Step 3:
[0156] The device transmits the stored sign language video data to the server via the internet.
[0157] Step 4:
[0158] The server analyzes the received sign language video data using a Convolutional Neural Network. It extracts characteristic patterns from the sign language and converts them into specific identifiers.
[0159] Step 5:
[0160] The server uses the identified identifier to query the database and retrieve background information and meanings for the corresponding sign language.
[0161] Step 6:
[0162] The server converts the background information and meaning retrieved from the database into structured data in JSON format and sends it to the terminal.
[0163] Step 7:
[0164] The structured data received by the terminal is displayed on the display device, and information such as "The sign language for 'thank you' is used to express gratitude" is presented to the user.
[0165] (Application Example 1)
[0166] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0167] Conventional sign language learning systems not only teach sign language but also lack systems that enable its practical application in real-world work environments. In particular, it is difficult for deaf individuals to give work instructions to machines and robots using sign language in factories and manufacturing sites. In such environments, a work instruction system that utilizes sign language is needed.
[0168] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0169] In this invention, the server includes means for analyzing sign language data, means for obtaining background information and meaning related to the analyzed sign language data from a database, and means for identifying work instructions based on the obtained background information and meaning and transmitting them to a device. This enables reliable transmission of work instructions using sign language.
[0170] A "user" refers to a person who uses the system to input sign language and issue work instructions.
[0171] "Sign language" refers to a system of gestures and hand movements used by people with hearing impairments to communicate.
[0172] "Means" refers to the methods, devices, and components used to realize each function of a system.
[0173] "Data" refers to information including video recordings of sign language and the results of their analysis that the system processes.
[0174] "Server" refers to a central processing unit that performs sign language data analysis and generates work instructions.
[0175] "Analysis" refers to the processing of an algorithm that receives sign language data, understands its content, and identifies its meaning.
[0176] "Background information" refers to the cultural and historical contextual information associated with the analyzed sign language.
[0177] "Meaning" refers to the content or intention that sign language is meant to convey.
[0178] A "database" refers to a collection of information that stores sign language, its background information, and its meaning in an associated manner.
[0179] A "terminal" refers to a device used by a user to input sign language and receive and display the analysis results.
[0180] "Equipment" refers to a device that receives and executes work instructions generated from analyzed sign language data.
[0181] "Work instructions" refer to specific operational instructions that robots and other automated equipment should perform based on the analyzed sign language.
[0182] "Execution" refers to the physical actions or processes that a device performs based on the work instructions it receives.
[0183] Modes for carrying out the invention
[0184] The system that realizes this invention involves a user inputting sign language, analyzing that data to generate relevant work instructions, and transmitting those instructions to a device for execution. This system is mainly composed of the following hardware and software.
[0185] hardware
[0186] Camera module: This utilizes cameras attached to factory robots, as well as the built-in cameras of smartphones and tablets. This allows for the capture of user sign language.
[0187] Factory robots: Specifically, this includes automation equipment from companies such as KUKA, ABB, and Fanuc. These robots are responsible for receiving and executing work instructions.
[0188] software
[0189] Image recognition algorithm: This algorithm analyzes sign language data using tools such as TensorFlow and OpenCV. It extracts features from sign language and identifies what those signs mean.
[0190] Sign language database: MongoDB or PostgreSQL is used. This database stores the meaning and background information of sign language.
[0191] Server-side program: Uses Node.js or Python (Flask) to analyze sign language data and generate work instructions.
[0192] Frontend: JavaScript® (React.js) is used to implement the user interface.
[0193] Specific examples of actions
[0194] 1. User inputs sign language: Users perform sign language towards a camera attached to the factory robot. Alternatively, they can input sign language from a smartphone or tablet. The sign language performed by the user mainly corresponds to work instructions such as "assemble" and "move".
[0195] Example prompt: "Perform the following sign language towards the camera and enter a work instruction: 'Move,' the robot will move."
[0196] 2. Data Transmission: The collected sign language data is sent to the server using WebSocket. The server divides the received sign language data into frames and prepares them for analysis.
[0197] 3. Data Analysis: An image recognition algorithm using TensorFlow is executed on the server side to analyze the received sign language data. This identifies the meaning of the sign language.
[0198] 4. Database query: The server uses the identified sign language identifier to retrieve the corresponding work instruction from a database such as MongoDB.
[0199] 5. Sending work instructions: The server sends the generated work instructions to the equipment. The equipment operates automatically based on the received work instructions.
[0200] Details of specific examples
[0201] For example, if a user performs the sign language for "move" while pointing it at the camera, the camera module captures the sign language. This data is sent to a server in real time and analyzed by TensorFlow. Based on the analysis results, the task instruction "move" is identified and retrieved from the database. The task instruction is then sent to the robot, which performs the specified movement.
[0202] This will enable people with hearing impairments to efficiently operate equipment in factories and manufacturing sites using sign language.
[0203] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0204] Step 1:
[0205] The user enters sign language.
[0206] The user performs sign language towards a camera attached to a factory robot or towards a smartphone or tablet. The entered sign language is captured as video and stored as digital data for use in the next processing step.
[0207] Input: User's sign language video
[0208] Output: Captured sign language video data
[0209] Step 2:
[0210] The device sends sign language data to the server.
[0211] The terminal (a robot or smartphone with a built-in camera module) divides the captured sign language video into frames in real time and sends them to the server using WebSocket.
[0212] Input: Captured sign language video data
[0213] Output: Sign language video frames sent to the server
[0214] Step 3:
[0215] The server analyzes the sign language data.
[0216] The server analyzes the received sign language data using a TensorFlow-based image recognition algorithm. It extracts each feature point of the sign language and performs analysis to determine what the sign language means.
[0217] Input: Sent sign language video frame
[0218] Output: Meaning of the analyzed sign language (identified sign language identifier)
[0219] Step 4:
[0220] The server retrieves background information and meaning from the database.
[0221] The server uses the identified sign language identifier to query a MongoDB or PostgreSQL database to retrieve background information, meaning, and corresponding work instructions for that sign language.
[0222] Input: Identified sign language identifier
[0223] Output: Background information, meaning, work instructions
[0224] Step 5:
[0225] The server sends work instructions to the equipment.
[0226] The server converts the acquired work instructions into structured data in JSON format and transmits it to equipment such as factory robots via the internet.
[0227] Input: Work instructions
[0228] Output: Work instructions sent to the device
[0229] Step 6:
[0230] The equipment performs tasks based on the work instructions it receives.
[0231] Factory automation equipment and robots analyze received work instructions and perform actions based on those instructions. For example, if the instruction is to move, the robot will move to the specified location.
[0232] Input: Received work instructions
[0233] Output: Performed actions (e.g., robot movement)
[0234] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0235] This invention relates to a system for learning sign language in a more accessible way, and in particular, a system that helps users input sign language and understand its background information, meaning, and even their emotions. The system consists of four main elements: the user, the terminal, the server, and the emotion engine.
[0236] Program Processing Overview
[0237] 1. The user enters the sign language.
[0238] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[0239] In this process, the device uses an emotion engine to acquire emotional data from the user's facial expressions and voice.
[0240] 2. The device sends sign language data and emotion data to the server.
[0241] The device stores sign language video captured by the camera and emotional data analyzed by the emotion engine as digital data.
[0242] Next, this sign language data and emotion data are sent to a server via the internet.
[0243] 3. The server analyzes the sign language data.
[0244] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Specifically, it extracts characteristic parts of the sign language and evaluates its similarity by comparing it with an existing sign language dictionary.
[0245] 4. The server identifies the sign language identifier.
[0246] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database.
[0247] 5. The server retrieves background information and meaning.
[0248] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[0249] 6. The server analyzes the emotional data.
[0250] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the user's emotional state is joyful when they input the sign language for "thank you."
[0251] 7. The server converts background information, semantic, and sentiment data into structured data.
[0252] The server converts the sign language background information, meaning, and emotional data into structured data and arranges it in a user-friendly format.
[0253] 8. The server sends the structured data to the terminal.
[0254] The server sends structured data regarding background information, meaning, and sentiment to the terminal.
[0255] 9. The device receives and displays information.
[0256] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, and videos, and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language.
[0257] Specific example
[0258] Example 1: Input of the sign language "thank you" and analysis of emotional data
[0259] The user performs the sign language for "thank you" while pointing it at the device's camera. The device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as the sign language for "thank you." It also uses an emotion engine to detect that the user's emotional state is joy. The server retrieves background information and meaning of the sign language for "thank you" from its database, adds the emotional state, and sends it to the device. The device then displays a message to the user stating something like, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0260] Example 2: Input of region-specific sign language and analysis of emotional data
[0261] The user inputs sign language specific to the Kansai region into the terminal. The terminal captures this sign language and sends it to the server along with emotion data. The server analyzes the sign language data and emotion data and recognizes that the sign language is specific to the Kansai region and that the user's emotional state is surprise or confusion. The server sends the regional background and cultural meaning of the sign language, as well as the emotional state, to the terminal. The terminal displays, "This sign language represents a greeting specific to the Kansai region, and the user's emotional state is 'surprise' or 'confusion'."
[0262] In this way, this system can improve the efficiency of sign language learning and the user's sense of familiarity by providing integrated background information, meaning, and the user's emotional state in sign language data.
[0263] The following describes the processing flow.
[0264] Step 1: The user enters sign language.
[0265] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[0266] The device captures sign language movements in real time using its camera.
[0267] Step 2: The device captures and saves sign language data and emotion data.
[0268] The device utilizes an emotion engine to analyze the user's emotional state in real time using facial recognition and voice analysis technologies.
[0269] These analysis results will be saved as digital data and managed together with the sign language data.
[0270] Step 3: The device sends sign language data and emotion data to the server.
[0271] The terminal transmits the stored sign language data and emotional data to the server via the internet. During this process, appropriate pre-processing, such as formatting and compression, is performed on the data.
[0272] Step 4: The server receives the sign language data.
[0273] The server receives sign language data and emotion data sent from the terminal.
[0274] Step 5: The server analyzes the sign language data.
[0275] The server analyzes the received sign language data and uses an image recognition algorithm to identify what the sign language gesture means. For example, it extracts the characteristic parts of the sign language gesture and compares them with an existing sign language dictionary.
[0276] Step 6: The server identifies the sign language identifier.
[0277] Based on the results of the image recognition algorithm, the server identifies the corresponding sign language identifier. This identifier corresponds to the sign language entry in the database.
[0278] Step 7: The server obtains background information and meaning.
[0279] Using the identified sign language identifier, the server obtains the background information (such as the region of occurrence, history, cultural background, etc.) and the meaning of the sign language from the database.
[0280] Step 8: The server analyzes the emotion data.
[0281] The server analyzes the received emotion data and evaluates the user's emotional state. For example, it determines whether the user is showing emotions such as "happiness" or "surprise".
[0282] Step 9: The server structures all the data.
[0283] The server integrates the background information, meaning, and emotion data of the sign language and structures them into a form that is easy for the user to understand.
[0284] Step 10: The server sends the structured data to the terminal.
[0285] The server sends the integrated structured data to the terminal.
[0286] Step 11: The terminal receives and displays the information.
[0287] The terminal receives structured data sent from the server and displays it to the user. The displayed content includes background information and meaning of sign language, as well as the user's emotional state. For example, it may be displayed visually in an easy-to-understand way using text, images, and videos.
[0288] Specific example
[0289] Example 1: Input of the sign language "thank you" and analysis of emotional data
[0290] 1. The user performs the sign language for "thank you" while pointing it at the device's camera.
[0291] 2. The device captures video and saves sign language data and emotion data analyzed by the emotion engine. The emotion engine detects the emotion of "joy" from the user's facial expressions and voice.
[0292] 3. The device sends sign language data and emotion data to the server.
[0293] 4. The server receives the sign language data.
[0294] 5. The server analyzes the sign language data using an image recognition algorithm.
[0295] 6. The server identifies the sign language identifier.
[0296] 7. The server retrieves background information and meaning of the sign language for "thank you" from the database.
[0297] 8. The server analyzes the emotion data and confirms that the user is experiencing the emotion of "joy."
[0298] 9. The server formats all data into structured data and sends it to the terminal.
[0299] 10. The device receives the data and displays the message, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0300] Example 2: Input of Region-Specific Sign Language and Analysis of Emotional Data
[0301] 1. The user inputs sign language specific to the Kansai region into the terminal.
[0302] 2. The terminal captures the sign language and transmits it to the server together with the emotional data. The emotion engine detects the emotion of "surprise" from the user's expression.
[0303] 3. The server analyzes the sign language data and emotional data to confirm that it is sign language specific to the region and the emotional state.
[0304] 4. The server obtains the regional background and cultural meaning from the database.
[0305] 5. The server integrates these data and transmits them to the terminal.
[0306] 6. The terminal displays "This sign language represents a greeting specific to the Kansai region, and the user's emotional state is'surprise'".
[0307] Through this system, the user can understand not only the background information and meaning of sign language but also their own emotional state, making sign language learning more accessible.
[0308] (Example 2)
[0309] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart device 14 is referred to as the "terminal".
[0310] In the conventional sign language learning system, there was a problem that it was impossible to simultaneously understand not only the background information and meaning of sign language but also the emotional state of the user performing the sign language. Also, although it has been required to make sign language learning more accessible and effective by reflecting the user's emotional state at the time of sign language input, there has been no effective system for this purpose.
[0311] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing transmitted sign language data and emotion data, means for obtaining background information and meaning related to the analyzed sign language data from a database, and means for evaluating the user's emotional state from the analyzed emotion data. This makes it possible to simultaneously grasp not only the background information and meaning of the sign language, but also the user's emotional state when performing the sign language.
[0312] A "user" refers to a person who uses the system to input sign language.
[0313] "Means of inputting sign language" refers to the interface (such as a camera or touchscreen) that a user uses to input sign language into a device.
[0314] "Sign language data" refers to information that digitally records the sign language actions entered by the user.
[0315] "Emotional data" refers to digital information that indicates the emotional state of a user, analyzed from their facial expressions, voice, and other data.
[0316] A "server" refers to a central computing device that performs data analysis, processing, and storage.
[0317] "Means of transmission" refers to the means of communication used to transfer data from a terminal to a server.
[0318] "Means of analysis" refers to algorithms and technologies (for example, image recognition algorithms and sentiment analysis technologies) that the server uses to process sign language data and sentiment data and analyze their content.
[0319] "Background information" refers to historical, cultural, and regional information related to sign language.
[0320] "Meaning" refers to the concept or intention expressed by the movements in sign language.
[0321] A "database" refers to an information management system that stores sign language data, background information, semantic data, and emotional data, making them accessible as needed.
[0322] A "specific identifier" refers to a code or label used to uniquely identify the type or meaning of sign language.
[0323] "Structured data" refers to data that follows a specific format and includes formalized data such as analyzed sign language information and emotional states.
[0324] A "terminal" refers to a device (such as a smartphone or tablet) used by a user to input sign language, receive data, and display it.
[0325] "Means of display" refers to the means by which a terminal visually presents analysis results and acquired information to the user.
[0326] This invention relates to a system for learning sign language in a more accessible way. The system helps users input sign language and understand its background information, meaning, and even their emotions. The system consists of four main elements: the user, the terminal, the server, and the emotion engine.
[0327] System Overview
[0328] 1. The user enters the sign language.
[0329] The user performs sign language actions in front of the device's camera. For example, to input the sign language for "thank you," the user lightly presses their palms in front of their chest. Alternatively, they can upload a pre-recorded video file of sign language to the device. In this case, the device's camera captures the user's sign language movements, and an emotion engine is used to obtain emotional data from the user's facial expressions and voice.
[0330] 2. The device sends sign language data and emotion data to the server.
[0331] The device captures sign language video with its camera and saves it as digital data. It also collects emotional data analyzed by an emotion engine. The captured sign language data and emotional data are sent together as a data packet to the server via the internet. Examples of emotion engines that can be used include "OpenCV" and "Microsoft® Azure® Face API".
[0332] 3. The server analyzes the sign language data.
[0333] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Examples of image recognition algorithms used include "TensorFlow" and "PyTorch". It extracts characteristic parts of the sign language and evaluates similarity by comparing them with existing sign language dictionaries.
[0334] 4. The server identifies the sign language identifier.
[0335] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database. For example, the sign language for "thank you" is identified as identifier "S001".
[0336] 5. The server retrieves background information and meaning.
[0337] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[0338] 6. The server analyzes the emotional data.
[0339] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the emotional state when a user inputs the sign language for "thank you" is "joy." Possible emotional analysis platforms used include "IBM Watson®" and "Amazon Rekognition."
[0340] 7. The server converts background information, semantic, and sentiment data into structured data.
[0341] The server converts the sign language background information, meaning, and emotional data into structured data and arranges it in a user-friendly format. For example, it converts it into a structured data format such as JSON.
[0342] 8. The server sends the structured data to the terminal.
[0343] The server sends structured data containing background information, semantic, and sentiment data to the terminal. Communication protocols used include HTTP and WebSocket.
[0344] 9. The device receives and displays information.
[0345] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, and videos, and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language. For example, the terminal screen may display the message, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0346] Specific example
[0347] Example 1: Input of the sign language "thank you" and analysis of emotional data
[0348] The user performs the sign language for "thank you" towards the device's camera, and the device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as identifier "S001". It also uses an emotion engine to detect that the user's emotional state is "joy". The server retrieves background information and meaning of the "thank you" sign language from the database, adds the emotional state, and sends it to the device. The device displays to the user, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0349] Example 2: Input of region-specific sign language and analysis of emotional data
[0350] The user inputs a sign language specific to the Kansai region into the terminal, and the terminal captures this sign language and sends it to the server along with emotion data. The server analyzes the sign language data and emotion data and recognizes that the sign language is specific to the Kansai region and that the user's emotional state is "surprise" or "confusion". The server sends the regional background and cultural meaning of the sign language, as well as the emotional state, to the terminal. The terminal displays, "This sign language expresses a greeting specific to the Kansai region, and the user's emotional state is 'surprise' or 'confusion'."
[0351] Example of a prompt
[0352] 1. "Please analyze the emotions expressed when performing sign language and provide this information along with the background context."
[0353] 2. What system can analyze the sign language for "thank you" and display the emotional state?
[0354] In this way, this system can improve the efficiency of sign language learning and the user's sense of familiarity by providing integrated background information, meaning, and the user's emotional state in sign language data.
[0355] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0356] Step 1:
[0357] The user enters sign language.
[0358] The user performs sign language actions in front of the device's camera. For example, to sign "thank you," the user lightly presses their palms in front of their chest. Alternatively, they can upload a pre-recorded sign language video file to the device. Input can be real-time sign language video or uploaded video files. Output is the device acquiring this sign language video data.
[0359] Step 2:
[0360] The device sends sign language data and emotion data to the server.
[0361] The device stores sign language video captured by its camera as digital data and uses an emotion engine to analyze emotional data from the user's facial expressions and voice. For example, the emotion engine uses facial recognition technology to analyze the degree of the user's smile and the tone of their voice. The input consists of the captured sign language video data and the results of the emotion analysis. As output, this data is sent to the server as data packets.
[0362] Step 3:
[0363] The server analyzes the sign language data.
[0364] The server inputs the received sign language data into an image recognition algorithm, extracts the characteristic parts of the sign language, and begins analysis. Examples of algorithms used include "TensorFlow" and "PyTorch." The input is sign language video data. The output is a sign language feature vector.
[0365] Step 4:
[0366] The server identifies the sign language identifier.
[0367] The server compares the generated feature vectors with existing sign language dictionaries to find the most similar sign language data. Each identified sign language data is assigned a unique identifier. The input consists of feature vectors and a sign language dictionary. The output is a sign language identifier (e.g., "S001").
[0368] Step 5:
[0369] The server retrieves background information and meaning.
[0370] The server uses the identified sign language identifier to retrieve background information (such as the region of origin, history, and cultural background) and its meaning from the database. The input is the sign language identifier. The output is the background information and meaning of the sign language.
[0371] Step 6:
[0372] The server analyzes the emotional data.
[0373] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the user's emotional state when they sign "thank you" is "joy." The input is emotional data, and the output is an evaluation of the user's emotional state.
[0374] Step 7:
[0375] The server converts background information, semantic, and sentiment data into structured data.
[0376] The server converts the sign language background information, meaning, and analyzed sentiment data into structured data (e.g., JSON format) and arranges it in a user-friendly format. Input includes background information, meaning, and sentiment data. The output is structured data.
[0377] Step 8:
[0378] The server sends structured data to the terminal.
[0379] The server sends structured data to the terminal. Communication protocols used include "HTTP" and "WebSocket". The input is structured data. The output is this data sent to the terminal.
[0380] Step 9:
[0381] The device receives and displays information.
[0382] The terminal receives structured data sent from the server and passes it to the UI for display to the user. The display format includes text, images, videos, etc., and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language. The input is structured data. The output is information that is displayed to the user. For example, the terminal screen might display "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0383] (Application Example 2)
[0384] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0385] This invention relates to a system for improving the ease of use and usability of electronic payment systems for users who use sign language. Conventional electronic payment systems are based on voice or text-based input, making them difficult to operate for users who use sign language. Furthermore, they lacked the function to analyze the user's emotional state and provide appropriate feedback accordingly. Therefore, there is a need to provide a system that integrates sign language input and emotion analysis to enable users who use sign language to comfortably make electronic payments.
[0386] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing input sign language data and user emotion data, means for obtaining background information and meaning related to the analyzed sign language data, as well as user emotion information, from a database, and means for transmitting the obtained background information, meaning, and emotion information to the terminal. As a result, when a user who uses sign language uses the electronic payment system, not only is the background information and meaning of the sign language displayed in an easy-to-understand manner, but appropriate feedback according to the user's emotional state is also provided.
[0387] A "user" refers to an individual who operates the system and inputs sign language and emotional data.
[0388] "Sign language data" refers to information that digitizes actions entered by users using sign language.
[0389] "Emotional data" refers to digitized information that represents the emotional state of a user, analyzed from their facial expressions and voice.
[0390] A "terminal" refers to a device used by users to input sign language and emotional data and send that information to a server.
[0391] A "server" refers to a device that analyzes sign language data and emotional data, retrieves the results from a database, and transmits them to a terminal.
[0392] An "emotion recognition algorithm" refers to a computational method for analyzing a user's emotional state from their facial expressions and voice.
[0393] An "image recognition algorithm" refers to a computational method for analyzing sign language data and identifying sign language movements.
[0394] An "identifier" refers to a specific feature or attribute associated with the analyzed sign language data and sentiment data.
[0395] "Background information" refers to information related to sign language, such as its place of origin, history, and cultural background.
[0396] "Meaning" refers to the concept or intention represented by a particular sign language action.
[0397] An "electronic payment system" refers to a system that allows users to conduct online transactions using digital currencies, credits, etc.
[0398] This invention provides a system that enables users who use sign language to smoothly utilize electronic payment systems. This system is implemented with the following configuration and means.
[0399] First, users input sign language using a device such as a smartphone. The device has a built-in camera, which can capture the user's sign language movements as video. Users can also input voice, which allows for the acquisition of facial expressions and vocal information. This data is collected as "sign language data" and "emotional data."
[0400] Next, the terminal sends the collected sign language data and emotion data to the server. An internet connection is used for transmission. The server implements image recognition algorithms and emotion recognition algorithms for analyzing the sign language and emotion data. Specifically, software such as TensorFlow and OpenCV are used. This analyzes the sign language movements, facial expressions, and voice to identify the user's intended meaning.
[0401] When the server analyzes sign language data and emotion data, it retrieves background information and meaning from a database based on the content. This database stores information such as the region of origin, history, and cultural background of the sign language, as well as the meaning of the sign language. The emotion data also includes the user's emotional state, which is also analyzed by the server. Based on the results of this analysis, the server retrieves background information and meaning of the sign language, as well as the user's emotional state, from the database along with a specific identifier.
[0402] Next, the server sends the acquired background information, meaning, and emotional information to the terminal. The terminal displays the received information to the user in an easy-to-understand format. The display format includes text, images, videos, etc., and is provided in a way that is easy for the user to understand. This ensures that when a user of sign language uses the electronic payment system, not only is the background information and meaning of the sign language clearly displayed, but appropriate feedback corresponding to the emotions the user is feeling is also provided at the same time.
[0403] As a concrete example, consider a case where a user performs the sign language for "thank you" in front of their smartphone camera. The device captures the video and sends it to the server. The server analyzes the sign language data to identify it as the sign language for "thank you," and uses an emotion recognition algorithm to detect that the user's emotional state is joy. The server retrieves background information and meaning of the "thank you" sign language, as well as the user's emotional state, from a database and sends it to the device. The device then displays a message to the user stating something like, "'Thank you' is a sign language expression used to express gratitude, and the user's emotional state is 'joy'."
[0404] Examples of prompt messages are as follows:
[0405] example:
[0406] Please generate a Python program that analyzes the following video and recognizes the meaning of the sign language and the user's emotions. The sign language should be returned as Unicode, and the emotions as text.
[0407] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0408] Step 1:
[0409] Users input sign language using their smartphones. The device's camera is used as the input device, capturing the user's sign language movements in real time. The user's facial expressions and voice are also recorded simultaneously. This data is collected as "sign language data" and "emotional data."
[0410] Input: User's sign language gestures, facial expressions, and voice.
[0411] Output: Sign language data, emotion data (captured video and audio)
[0412] Step 2:
[0413] The terminal collects sign language data and emotional data and sends it to the server. An internet connection is used for transmission. Furthermore, the data undergoes necessary preprocessing (e.g., frame resizing and normalization) before being sent to the server.
[0414] Input: Sign language data, emotion data
[0415] Output: Digital data sent to the server
[0416] Step 3:
[0417] The server analyzes sign language data and emotion data. Sign language data is analyzed using an image recognition algorithm (e.g., a TensorFlow model) to identify the gestures and meanings of the signs. Emotion data is analyzed using an emotion recognition algorithm (e.g., a TensorFlow / Keras model) to identify the user's emotional state.
[0418] Input: Sign language data, emotion data
[0419] Output: Analysis results (sign language identification information, emotional state)
[0420] Step 4:
[0421] Based on the analysis results, the server retrieves background information and meanings of sign language, as well as user emotional information, from the database. The database stores the region of origin, history, cultural background, and meaning of each sign language. Appropriate feedback information tailored to the user's emotional state is also retrieved.
[0422] Input: Analysis results (sign language identification information, emotional state)
[0423] Output: Sign language background information, semantic information, and emotional information
[0424] Step 5:
[0425] The server transmits background information, semantic information, and emotional information acquired by the server to the terminal. The data is structured and presented in a user-friendly format. An internet connection is used for transmission.
[0426] Input: Background information, meaning, sentiment information
[0427] Output: Structured data sent to the terminal
[0428] Step 6:
[0429] The device displays received information to the user. The display format includes text, images, and videos, and is provided in a way that allows the user to easily understand the background information, meaning, and emotional state of the sign language.
[0430] Input: Received structured data
[0431] Output: Information displayed to the user (background information, meaning, emotional state)
[0432] Through these steps, users can smoothly utilize the electronic payment system using sign language, and in the process, receive appropriate feedback that reflects the background information and meaning of the sign language, as well as their own emotional state.
[0433] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0434] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0435] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0436] [Second Embodiment]
[0437] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0438] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0439] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0440] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0441] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0442] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0443] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0444] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0445] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0446] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0447] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0448] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0449] This invention relates to a system for learning sign language in a more accessible way, and in particular, a system that helps users input sign language and understand its background information and meaning. The system consists of three main elements: the user, the terminal, and the server.
[0450] Program Processing Overview
[0451] 1. The user enters the sign language.
[0452] Users perform sign language using their device's camera. Alternatively, they can upload pre-recorded sign language video files to their device.
[0453] 2. The device sends sign language data to the server.
[0454] The device saves the sign language video captured by its camera as digital data. Next, it transmits this sign language data to a server via the internet.
[0455] 3. The server analyzes the sign language data.
[0456] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Specifically, it extracts characteristic parts of the sign language and compares them with an existing sign language dictionary.
[0457] 4. The server retrieves background information and meaning.
[0458] The server queries the database using the identified sign language identifier to retrieve background information and meanings for that sign language. For example, it may retrieve information about the origins of the sign language, as well as its cultural and historical background.
[0459] 5. The server sends information to the terminal.
[0460] The server converts the acquired background information and its meaning into structured data and sends it to the terminal.
[0461] 6. The device displays the information.
[0462] The terminal receives background information and meaning from the server and displays it to the user. Display formats include text, images, and videos.
[0463] Specific example
[0464] Example 1: Inputting the sign language for "thank you"
[0465] The user performs the sign language for "thank you" while pointing it at the device's camera. The device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as the sign language for "thank you." The server retrieves background information and meaning of the sign language for "thank you" from its database and sends it to the device. The device then displays this information to the user. Specifically, it might display something like, "The sign language for 'thank you' is used to express gratitude in Japanese."
[0466] Example 2: Inputting regional sign language
[0467] The user inputs sign language specific to the Kansai region into the terminal. The terminal captures this sign language and sends it to the server. The server analyzes the sign language data and recognizes that this sign language is specific to the Kansai region. The server retrieves the regional background and cultural meaning of this sign language from its database and sends it to the terminal. The terminal displays this information to the user. Specifically, it displays, "This sign language represents a greeting unique to people in the Kansai region."
[0468] In this way, by providing users with background information and meanings of sign language data, this system can improve the efficiency of sign language learning and enhance users' familiarity with sign language.
[0469] The following describes the processing flow.
[0470] Step 1: The user enters sign language.
[0471] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[0472] Step 2: The device captures and saves the sign language data.
[0473] The device saves sign language videos captured by the camera as digital data. Similarly, uploaded video files are also saved as digital data.
[0474] Step 3: The device sends the sign language data to the server.
[0475] The terminal transmits the stored sign language data to the server via the internet. During this process, appropriate pre-processing, such as formatting and compression, is performed on the data.
[0476] Step 4: The server receives the sign language data.
[0477] The server receives sign language data sent from the terminal.
[0478] Step 5: The server analyzes the sign language data.
[0479] The server applies an image recognition algorithm to analyze the received sign language data. It extracts characteristic parts of sign language movements and evaluates their similarity by comparing them with an existing sign language dictionary.
[0480] Step 6: The server identifies the sign language identifier.
[0481] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database.
[0482] Step 7: The server retrieves background information and meaning.
[0483] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[0484] Step 8: The server converts background information and meaning into structured data.
[0485] The server converts the acquired background information and meaning into structured data in a format that is easy for the user to read.
[0486] Step 9: The server sends the structured data to the terminal.
[0487] The server sends structured data containing background information and semantics to the terminal.
[0488] Step 10: The device receives and displays information.
[0489] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, videos, etc., and is provided in a way that makes it easy for the user to understand the background information and meaning of sign language.
[0490] In this way, users can understand detailed background information and meaning related to the sign language they input. This process is designed to make learning sign language efficient and accessible.
[0491] (Example 1)
[0492] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0493] When learning sign language, many users face the challenge of understanding the precise meaning and background information of sign language movements. Furthermore, existing sign language learning systems often suffer from low accuracy in analyzing sign language or insufficient acquisition of background information, which hinders the improvement of sign language learning efficiency.
[0494] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0495] In this invention, the server includes means for analyzing transmitted sign language data using a machine learning algorithm, means for obtaining background information and meaning related to the analyzed sign language data from a data store, and means for sending the obtained background information and meaning to a terminal. As a result, when a user inputs sign language, they can quickly obtain the accurate meaning and background information of that sign language.
[0496] A "user" is a person who uses this system to input sign language and receives the analysis results and background information.
[0497] An "input device" is a device used by a user to input sign language, and specifically refers to a camera or touchscreen.
[0498] A "communication device" refers to a device used to transmit input sign language data to a server, and specifically includes wireless communication modules and internet connectivity functions.
[0499] A "server" is a computing device responsible for analyzing sign language data, acquiring background information and meaning, and transmitting this information to users.
[0500] A "machine learning algorithm" is a mathematical method used to analyze sign language data and identify its meaning, specifically referring to Convolutional Neural Networks (CNNs), among others.
[0501] A "data store" refers to a database or storage system used to store background information and meanings of sign language.
[0502] A "display device" is a device used to visually display received information to a user, and specifically refers to displays and monitors.
[0503] A "unique identifier" is a code or number used to uniquely identify a particular sign language, and is used in database queries.
[0504] Modes for carrying out the invention
[0505] This invention relates to a system that helps users input sign language and understand its background information and meaning. The following describes specific embodiments for carrying out the invention.
[0506] Hardware and software to be used
[0507] 1. User
[0508] The user's role is to input sign language. This is done using input devices equipped with cameras and touchscreens.
[0509] 2. Terminal
[0510] The terminal's role is to capture sign language data entered by the user and send it to the server. The terminal's communication equipment includes a wireless communication module and internet connectivity.
[0511] The terminal includes video capture devices (e.g., cameras) and data display devices (e.g., displays).
[0512] 3. Server
[0513] The server's role is to analyze sign language data, extract background information and meaning, and provide it to the user. Sign language analysis utilizes image recognition algorithms such as Convolutional Neural Networks (CNNs). Specifically, machine learning frameworks like TensorFlow and PyTorch are used.
[0514] The server also has a data transmission function to send the acquired background information and meaning to the terminal.
[0515] 4. Datastore
[0516] Data stores are used to store background information and meanings of sign language. SQL and NoSQL databases are typical examples.
[0517] Specific example
[0518] Example 1: Inputting "thank you" in sign language
[0519] 1. The user performs the sign language for "thank you" while pointing it at the device's camera.
[0520] The user performs the sign language for "thank you" in front of the camera.
[0521] 2. The device captures this video and sends it to the server as sign language data.
[0522] The device's camera captures the video frame by frame and saves it as digital data. The data is then sent to a server.
[0523] 3. The server analyzes the sign language data and identifies it as the sign for "thank you."
[0524] The server analyzes the received data using an image recognition algorithm and identifies it as the sign language for "thank you."
[0525] 4. The server retrieves background information and meaning of the sign language for "thank you" from the database and displays the retrieved information on the terminal.
[0526] The server retrieves information related to the sign language expression for "thank you" from its database. For example, this information may include details about how gratitude is expressed in Japanese culture. The retrieved information is then sent to the terminal as structured data (e.g., in JSON format).
[0527] 5. Display the information received by the device to the user.
[0528] The device visually displays the information it receives to the user. For example, it might display information such as, "The sign language for 'thank you' is used to express gratitude in Japanese."
[0529] Example of a prompt
[0530] The following is an example of a prompt statement for inputting a system description into the generating AI model.
[0531] Prompt message:
[0532] I would like to describe a system for learning sign language. This system helps users input sign language and understand its meaning and background information. Please explain the program's process in natural language, following the steps below.
[0533] procedure:
[0534] 1. The user inputs the sign language. The user either performs the sign language using their device's camera or uploads a pre-recorded video file of the sign language.
[0535] 2. The device sends sign language data to the server. The device saves the sign language video captured by the camera as digital data and sends it to the server via the internet.
[0536] 3. The server analyzes the sign language data. The server uses an image recognition algorithm to analyze what the sign language data means and compares it to an existing sign language dictionary.
[0537] 4. The server retrieves background information and meaning. The server queries the database using the identified sign language identifier to retrieve background information and meaning for that sign language.
[0538] 5. The server sends the information to the terminal. The server converts the acquired background information and meaning into structured data and sends it to the terminal.
[0539] 6. The terminal displays information. The terminal receives background information and meaning sent from the server and displays it to the user. Display formats include text, images, videos, etc.
[0540] Specific example:
[0541] Please explain the processing steps using the sign language for "thank you" as an example.
[0542] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0543] Specific flow of program processing
[0544] Processing steps
[0545] Step 1:
[0546] The user inputs sign language. The user either performs sign language using the camera on their device or uploads a pre-recorded video file of sign language. The input data is a video of sign language. This video of sign language becomes the raw data used for analysis in subsequent processing.
[0547] Step 2:
[0548] The device captures sign language data and saves it as digital data. The device's camera captures video frame by frame and saves this as digital data (e.g., MP4 format) to internal storage. The captured data is used for subsequent analysis.
[0549] Step 3:
[0550] The terminal transmits the stored sign language data to the server via a communication device. The terminal uses an internet connection to transmit the stored sign language data to the server. The input data is the captured sign language video data, and the output is the data transmitted to the server.
[0551] Step 4:
[0552] The server receives sign language data and analyzes it using a machine learning algorithm. The server then analyzes the received data using an image recognition algorithm (e.g., a Convolutional Neural Network). This analysis extracts the features of the sign language and converts them into an identifiable format. The input data is video data of sign language, and the output is the analyzed sign language data.
[0553] Step 5:
[0554] The server identifies a unique identifier based on the analyzed sign language data and retrieves background information and meaning from the data store. The server queries the database based on the analysis results to retrieve background information and meaning for the corresponding sign language. The input data is the analyzed sign language data, and the output is data related to background information and meaning.
[0555] Step 6:
[0556] The server converts the acquired information into structured data (e.g., JSON format) and sends it to the terminal. The server uses a data transmission function to send the acquired information to the terminal. The input data is background information and semantic data, and the output is the structured data sent to the terminal.
[0557] Step 7:
[0558] The terminal displays the information it receives to the user. The terminal's display device analyzes the received structured data and displays it visually to the user. Display formats include text, images, and videos. The input data is structured data, and the output is the information displayed to the user.
[0559] Detailed explanation of operation
[0560] Step 1:
[0561] Users perform the sign language for "thank you" using their camera. Users can also upload the recorded video file.
[0562] Step 2:
[0563] The device's camera captures video of the user performing sign language. This video is saved as digital data to the internal storage.
[0564] Step 3:
[0565] The device transmits the stored sign language video data to the server via the internet.
[0566] Step 4:
[0567] The server analyzes the received sign language video data using a Convolutional Neural Network. It extracts characteristic patterns from the sign language and converts them into specific identifiers.
[0568] Step 5:
[0569] The server uses the identified identifier to query the database and retrieve background information and meanings for the corresponding sign language.
[0570] Step 6:
[0571] The server converts the background information and meaning retrieved from the database into structured data in JSON format and sends it to the terminal.
[0572] Step 7:
[0573] The structured data received by the terminal is displayed on the display device, and information such as "The sign language for 'thank you' is used to express gratitude" is presented to the user.
[0574] (Application Example 1)
[0575] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0576] Conventional sign language learning systems not only teach sign language but also lack systems that enable its practical application in real-world work environments. In particular, it is difficult for deaf individuals to give work instructions to machines and robots using sign language in factories and manufacturing sites. In such environments, a work instruction system that utilizes sign language is needed.
[0577] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0578] In this invention, the server includes means for analyzing sign language data, means for obtaining background information and meaning related to the analyzed sign language data from a database, and means for identifying work instructions based on the obtained background information and meaning and transmitting them to a device. This enables reliable transmission of work instructions using sign language.
[0579] A "user" refers to a person who uses the system to input sign language and issue work instructions.
[0580] "Sign language" refers to a system of gestures and hand movements used by people with hearing impairments to communicate.
[0581] "Means" refers to the methods, devices, and components used to realize each function of a system.
[0582] "Data" refers to information including video recordings of sign language and the results of their analysis that the system processes.
[0583] "Server" refers to a central processing unit that performs sign language data analysis and generates work instructions.
[0584] "Analysis" refers to the processing of an algorithm that receives sign language data, understands its content, and identifies its meaning.
[0585] "Background information" refers to the cultural and historical contextual information associated with the analyzed sign language.
[0586] "Meaning" refers to the content or intention that sign language is meant to convey.
[0587] A "database" refers to a collection of information that stores sign language, its background information, and its meaning in an associated manner.
[0588] A "terminal" refers to a device used by a user to input sign language and receive and display the analysis results.
[0589] "Equipment" refers to a device that receives and executes work instructions generated from analyzed sign language data.
[0590] "Work instructions" refer to specific operational instructions that robots and other automated equipment should perform based on the analyzed sign language.
[0591] "Execution" refers to the physical actions or processes that a device performs based on the work instructions it receives.
[0592] Modes for carrying out the invention
[0593] The system that realizes this invention involves a user inputting sign language, analyzing that data to generate relevant work instructions, and transmitting those instructions to a device for execution. This system is mainly composed of the following hardware and software.
[0594] hardware
[0595] Camera module: This utilizes cameras attached to factory robots, as well as the built-in cameras of smartphones and tablets. This allows for the capture of user sign language.
[0596] Factory robots: Specifically, this includes automation equipment from companies such as KUKA, ABB, and Fanuc. These robots are responsible for receiving and executing work instructions.
[0597] software
[0598] Image recognition algorithm: This algorithm analyzes sign language data using tools such as TensorFlow and OpenCV. It extracts features from sign language and identifies what those signs mean.
[0599] Sign language database: MongoDB or PostgreSQL is used. This database stores the meaning and background information of sign language.
[0600] Server-side program: Uses Node.js or Python (Flask) to analyze sign language data and generate work instructions.
[0601] Frontend: JavaScript (React.js) is used to implement the user interface.
[0602] Specific examples of actions
[0603] 1. User inputs sign language: Users perform sign language towards a camera attached to the factory robot. Alternatively, they can input sign language from a smartphone or tablet. The sign language performed by the user mainly corresponds to work instructions such as "assemble" and "move".
[0604] Example prompt: "Perform the following sign language towards the camera and enter a work instruction: 'Move,' the robot will move."
[0605] 2. Data Transmission: The collected sign language data is sent to the server using WebSocket. The server divides the received sign language data into frames and prepares them for analysis.
[0606] 3. Data Analysis: An image recognition algorithm using TensorFlow is executed on the server side to analyze the received sign language data. This identifies the meaning of the sign language.
[0607] 4. Database query: The server uses the identified sign language identifier to retrieve the corresponding work instruction from a database such as MongoDB.
[0608] 5. Sending work instructions: The server sends the generated work instructions to the equipment. The equipment operates automatically based on the received work instructions.
[0609] Details of specific examples
[0610] For example, if a user performs the sign language for "move" while pointing it at the camera, the camera module captures the sign language. This data is sent to a server in real time and analyzed by TensorFlow. Based on the analysis results, the task instruction "move" is identified and retrieved from the database. The task instruction is then sent to the robot, which performs the specified movement.
[0611] This will enable people with hearing impairments to efficiently operate equipment in factories and manufacturing sites using sign language.
[0612] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0613] Step 1:
[0614] The user enters sign language.
[0615] The user performs sign language towards a camera attached to a factory robot or towards a smartphone or tablet. The entered sign language is captured as video and stored as digital data for use in the next processing step.
[0616] Input: User's sign language video
[0617] Output: Captured sign language video data
[0618] Step 2:
[0619] The device sends sign language data to the server.
[0620] The terminal (a robot or smartphone with a built-in camera module) divides the captured sign language video into frames in real time and sends them to the server using WebSocket.
[0621] Input: Captured sign language video data
[0622] Output: Sign language video frames sent to the server
[0623] Step 3:
[0624] The server analyzes the sign language data.
[0625] The server analyzes the received sign language data using a TensorFlow-based image recognition algorithm. It extracts each feature point of the sign language and performs analysis to determine what the sign language means.
[0626] Input: Sent sign language video frame
[0627] Output: Meaning of the analyzed sign language (identified sign language identifier)
[0628] Step 4:
[0629] The server retrieves background information and meaning from the database.
[0630] The server uses the identified sign language identifier to query a MongoDB or PostgreSQL database to retrieve background information, meaning, and corresponding work instructions for that sign language.
[0631] Input: Identified sign language identifier
[0632] Output: Background information, meaning, work instructions
[0633] Step 5:
[0634] The server sends work instructions to the equipment.
[0635] The server converts the acquired work instructions into structured data in JSON format and transmits it to equipment such as factory robots via the internet.
[0636] Input: Work instructions
[0637] Output: Work instructions sent to the device
[0638] Step 6:
[0639] The equipment performs tasks based on the work instructions it receives.
[0640] Factory automation equipment and robots analyze received work instructions and perform actions based on those instructions. For example, if the instruction is to move, the robot will move to the specified location.
[0641] Input: Received work instructions
[0642] Output: Performed actions (e.g., robot movement)
[0643] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0644] This invention relates to a system for learning sign language in a more accessible way, and in particular, a system that helps users input sign language and understand its background information, meaning, and even their emotions. The system consists of four main elements: the user, the terminal, the server, and the emotion engine.
[0645] Program Processing Overview
[0646] 1. The user enters the sign language.
[0647] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[0648] In this process, the device uses an emotion engine to acquire emotional data from the user's facial expressions and voice.
[0649] 2. The device sends sign language data and emotion data to the server.
[0650] The device stores sign language video captured by the camera and emotional data analyzed by the emotion engine as digital data.
[0651] Next, this sign language data and emotion data are sent to a server via the internet.
[0652] 3. The server analyzes the sign language data.
[0653] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Specifically, it extracts characteristic parts of the sign language and evaluates its similarity by comparing it with an existing sign language dictionary.
[0654] 4. The server identifies the sign language identifier.
[0655] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database.
[0656] 5. The server retrieves background information and meaning.
[0657] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[0658] 6. The server analyzes the emotional data.
[0659] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the user's emotional state is joyful when they input the sign language for "thank you."
[0660] 7. The server converts background information, semantic, and sentiment data into structured data.
[0661] The server converts the sign language background information, meaning, and emotional data into structured data and arranges it in a user-friendly format.
[0662] 8. The server sends the structured data to the terminal.
[0663] The server sends structured data regarding background information, meaning, and sentiment to the terminal.
[0664] 9. The device receives and displays information.
[0665] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, and videos, and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language.
[0666] Specific example
[0667] Example 1: Input of the sign language "thank you" and analysis of emotional data
[0668] The user performs the sign language for "thank you" while pointing it at the device's camera. The device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as the sign language for "thank you." It also uses an emotion engine to detect that the user's emotional state is joy. The server retrieves background information and meaning of the sign language for "thank you" from its database, adds the emotional state, and sends it to the device. The device then displays a message to the user stating something like, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0669] Example 2: Input of region-specific sign language and analysis of emotional data
[0670] The user inputs sign language specific to the Kansai region into the terminal. The terminal captures this sign language and sends it to the server along with emotion data. The server analyzes the sign language data and emotion data and recognizes that the sign language is specific to the Kansai region and that the user's emotional state is surprise or confusion. The server sends the regional background and cultural meaning of the sign language, as well as the emotional state, to the terminal. The terminal displays, "This sign language represents a greeting specific to the Kansai region, and the user's emotional state is 'surprise' or 'confusion'."
[0671] In this way, this system can improve the efficiency of sign language learning and the user's sense of familiarity by providing integrated background information, meaning, and the user's emotional state in sign language data.
[0672] The following describes the processing flow.
[0673] Step 1: The user enters sign language.
[0674] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[0675] The device captures sign language movements in real time using its camera.
[0676] Step 2: The device captures and saves sign language data and emotion data.
[0677] The device utilizes an emotion engine to analyze the user's emotional state in real time using facial recognition and voice analysis technologies.
[0678] These analysis results will be saved as digital data and managed together with the sign language data.
[0679] Step 3: The device sends sign language data and emotion data to the server.
[0680] The terminal transmits the stored sign language data and emotional data to the server via the internet. During this process, appropriate pre-processing, such as formatting and compression, is performed on the data.
[0681] Step 4: The server receives the sign language data.
[0682] The server receives sign language data and emotion data sent from the terminal.
[0683] Step 5: The server analyzes the sign language data.
[0684] The server analyzes the received sign language data and uses image recognition algorithms to identify what the sign language actions mean. For example, it extracts characteristic parts of the sign language actions and compares them with existing sign language dictionaries.
[0685] Step 6: The server identifies the sign language identifier.
[0686] The server identifies the corresponding sign language identifier based on the results of the image recognition algorithm. This identifier corresponds to a sign language entry in the database.
[0687] Step 7: The server retrieves background information and meaning.
[0688] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[0689] Step 8: The server analyzes the emotion data.
[0690] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it determines whether the user is exhibiting emotions such as "joy" or "surprise."
[0691] Step 9: The server structures all the data.
[0692] The server integrates background information, meaning, and emotional data from sign language and structures it into a format that is easy for users to understand.
[0693] Step 10: The server sends the structured data to the terminal.
[0694] The server sends integrated structured data to the terminal.
[0695] Step 11: The device receives and displays information.
[0696] The terminal receives structured data sent from the server and displays it to the user. The displayed content includes background information and meaning of sign language, as well as the user's emotional state. For example, it may be displayed visually in an easy-to-understand way using text, images, and videos.
[0697] Specific example
[0698] Example 1: Input of the sign language "thank you" and analysis of emotional data
[0699] 1. The user performs the sign language for "thank you" while pointing it at the device's camera.
[0700] 2. The device captures video and saves sign language data and emotion data analyzed by the emotion engine. The emotion engine detects the emotion of "joy" from the user's facial expressions and voice.
[0701] 3. The device sends sign language data and emotion data to the server.
[0702] 4. The server receives the sign language data.
[0703] 5. The server analyzes the sign language data using an image recognition algorithm.
[0704] 6. The server identifies the sign language identifier.
[0705] 7. The server retrieves background information and meaning of the sign language for "thank you" from the database.
[0706] 8. The server analyzes the emotion data and confirms that the user is experiencing the emotion of "joy."
[0707] 9. The server formats all data into structured data and sends it to the terminal.
[0708] 10. The device receives the data and displays the message, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0709] Example 2: Input of region-specific sign language and analysis of emotional data
[0710] 1. The user inputs sign language specific to the Kansai region into the device.
[0711] 2. The device captures sign language and sends it to the server along with emotion data. The emotion engine detects the emotion of "surprise" from the user's facial expressions.
[0712] 3. The server analyzes the sign language data and emotion data to confirm that it is a regionally specific sign language and to verify the emotional state.
[0713] 4. The server retrieves regional background and cultural significance from the database.
[0714] 5. The server integrates this data and sends it to the terminal.
[0715] 6. The device displays the message, "This sign language represents a greeting unique to the Kansai region, and the user's emotional state is 'surprise'."
[0716] Through this system, users can understand not only the background information and meaning of sign language, but also their own emotional state, making sign language learning more accessible.
[0717] (Example 2)
[0718] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0719] Conventional sign language learning systems have the drawback of not being able to simultaneously understand not only the background information and meaning of sign language, but also the emotional state of the user performing the sign language. Furthermore, there is a need to make learning sign language more accessible and effective by reflecting the user's emotional state when inputting sign language, but an effective system for this purpose has not existed.
[0720] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing transmitted sign language data and emotion data, means for obtaining background information and meaning related to the analyzed sign language data from a database, and means for evaluating the user's emotional state from the analyzed emotion data. This makes it possible to simultaneously grasp not only the background information and meaning of the sign language, but also the user's emotional state when performing the sign language.
[0721] A "user" refers to a person who uses the system to input sign language.
[0722] "Means of inputting sign language" refers to the interface (such as a camera or touchscreen) that a user uses to input sign language into a device.
[0723] "Sign language data" refers to information that digitally records the sign language actions entered by the user.
[0724] "Emotional data" refers to digital information that indicates the emotional state of a user, analyzed from their facial expressions, voice, and other data.
[0725] A "server" refers to a central computing device that performs data analysis, processing, and storage.
[0726] "Means of transmission" refers to the means of communication used to transfer data from a terminal to a server.
[0727] "Means of analysis" refers to algorithms and technologies (for example, image recognition algorithms and sentiment analysis technologies) that the server uses to process sign language data and sentiment data and analyze their content.
[0728] "Background information" refers to historical, cultural, and regional information related to sign language.
[0729] "Meaning" refers to the concept or intention expressed by the movements in sign language.
[0730] A "database" refers to an information management system that stores sign language data, background information, semantic data, and emotional data, making them accessible as needed.
[0731] A "specific identifier" refers to a code or label used to uniquely identify the type or meaning of sign language.
[0732] "Structured data" refers to data that follows a specific format and includes formalized data such as analyzed sign language information and emotional states.
[0733] A "terminal" refers to a device (such as a smartphone or tablet) used by a user to input sign language, receive data, and display it.
[0734] "Means of display" refers to the means by which a terminal visually presents analysis results and acquired information to the user.
[0735] This invention relates to a system for learning sign language in a more accessible way. The system helps users input sign language and understand its background information, meaning, and even their emotions. The system consists of four main elements: the user, the terminal, the server, and the emotion engine.
[0736] System Overview
[0737] 1. The user enters the sign language.
[0738] The user performs sign language actions in front of the device's camera. For example, to input the sign language for "thank you," the user lightly presses their palms in front of their chest. Alternatively, they can upload a pre-recorded video file of sign language to the device. In this case, the device's camera captures the user's sign language movements, and an emotion engine is used to obtain emotional data from the user's facial expressions and voice.
[0739] 2. The device sends sign language data and emotion data to the server.
[0740] The device captures sign language video with its camera and saves it as digital data. It also collects emotional data analyzed by an emotion engine. The captured sign language data and emotional data are sent together as a data packet to the server via the internet. Examples of emotion engines that can be used include "OpenCV" and "Microsoft Azure Face API".
[0741] 3. The server analyzes the sign language data.
[0742] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Examples of image recognition algorithms used include "TensorFlow" and "PyTorch". It extracts characteristic parts of the sign language and evaluates similarity by comparing them with existing sign language dictionaries.
[0743] 4. The server identifies the sign language identifier.
[0744] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database. For example, the sign language for "thank you" is identified as identifier "S001".
[0745] 5. The server retrieves background information and meaning.
[0746] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[0747] 6. The server analyzes the emotional data.
[0748] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the emotional state of a user when they input the sign language for "thank you" is "joy." Possible emotional analysis platforms used include "IBM Watson" and "Amazon Rekognition."
[0749] 7. The server converts background information, semantic, and sentiment data into structured data.
[0750] The server converts the sign language background information, meaning, and emotional data into structured data and arranges it in a user-friendly format. For example, it converts it into a structured data format such as JSON.
[0751] 8. The server sends the structured data to the terminal.
[0752] The server sends structured data containing background information, semantic, and sentiment data to the terminal. Communication protocols used include HTTP and WebSocket.
[0753] 9. The device receives and displays information.
[0754] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, and videos, and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language. For example, the terminal screen may display the message, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0755] Specific example
[0756] Example 1: Input of the sign language "thank you" and analysis of emotional data
[0757] The user performs the sign language for "thank you" towards the device's camera, and the device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as identifier "S001". It also uses an emotion engine to detect that the user's emotional state is "joy". The server retrieves background information and meaning of the "thank you" sign language from the database, adds the emotional state, and sends it to the device. The device displays to the user, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0758] Example 2: Input of region-specific sign language and analysis of emotional data
[0759] The user inputs a sign language specific to the Kansai region into the terminal, and the terminal captures this sign language and sends it to the server along with emotion data. The server analyzes the sign language data and emotion data and recognizes that the sign language is specific to the Kansai region and that the user's emotional state is "surprise" or "confusion". The server sends the regional background and cultural meaning of the sign language, as well as the emotional state, to the terminal. The terminal displays, "This sign language expresses a greeting specific to the Kansai region, and the user's emotional state is 'surprise' or 'confusion'."
[0760] Example of a prompt
[0761] 1. "Please analyze the emotions expressed when performing sign language and provide this information along with the background context."
[0762] 2. What system can analyze the sign language for "thank you" and display the emotional state?
[0763] In this way, this system can improve the efficiency of sign language learning and the user's sense of familiarity by providing integrated background information, meaning, and the user's emotional state in sign language data.
[0764] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0765] Step 1:
[0766] The user enters sign language.
[0767] The user performs sign language actions in front of the device's camera. For example, to sign "thank you," the user lightly presses their palms in front of their chest. Alternatively, they can upload a pre-recorded sign language video file to the device. Input can be real-time sign language video or uploaded video files. Output is the device acquiring this sign language video data.
[0768] Step 2:
[0769] The device sends sign language data and emotion data to the server.
[0770] The device stores sign language video captured by its camera as digital data and uses an emotion engine to analyze emotional data from the user's facial expressions and voice. For example, the emotion engine uses facial recognition technology to analyze the degree of the user's smile and the tone of their voice. The input consists of the captured sign language video data and the results of the emotion analysis. As output, this data is sent to the server as data packets.
[0771] Step 3:
[0772] The server analyzes the sign language data.
[0773] The server inputs the received sign language data into an image recognition algorithm, extracts the characteristic parts of the sign language, and begins analysis. Examples of algorithms used include "TensorFlow" and "PyTorch." The input is sign language video data. The output is a sign language feature vector.
[0774] Step 4:
[0775] The server identifies the sign language identifier.
[0776] The server compares the generated feature vectors with existing sign language dictionaries to find the most similar sign language data. Each identified sign language data is assigned a unique identifier. The input consists of feature vectors and a sign language dictionary. The output is a sign language identifier (e.g., "S001").
[0777] Step 5:
[0778] The server retrieves background information and meaning.
[0779] The server uses the identified sign language identifier to retrieve background information (such as the region of origin, history, and cultural background) and its meaning from the database. The input is the sign language identifier. The output is the background information and meaning of the sign language.
[0780] Step 6:
[0781] The server analyzes the emotional data.
[0782] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the user's emotional state when they sign "thank you" is "joy." The input is emotional data, and the output is an evaluation of the user's emotional state.
[0783] Step 7:
[0784] The server converts background information, semantic, and sentiment data into structured data.
[0785] The server converts the sign language background information, meaning, and analyzed sentiment data into structured data (e.g., JSON format) and arranges it in a user-friendly format. Input includes background information, meaning, and sentiment data. The output is structured data.
[0786] Step 8:
[0787] The server sends structured data to the terminal.
[0788] The server sends structured data to the terminal. Communication protocols used include "HTTP" and "WebSocket". The input is structured data. The output is this data sent to the terminal.
[0789] Step 9:
[0790] The device receives and displays information.
[0791] The terminal receives structured data sent from the server and passes it to the UI for display to the user. The display format includes text, images, videos, etc., and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language. The input is structured data. The output is information that is displayed to the user. For example, the terminal screen might display "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[0792] (Application Example 2)
[0793] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0794] This invention relates to a system for improving the ease of use and usability of electronic payment systems for users who use sign language. Conventional electronic payment systems are based on voice or text-based input, making them difficult to operate for users who use sign language. Furthermore, they lacked the function to analyze the user's emotional state and provide appropriate feedback accordingly. Therefore, there is a need to provide a system that integrates sign language input and emotion analysis to enable users who use sign language to comfortably make electronic payments.
[0795] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing input sign language data and user emotion data, means for obtaining background information and meaning related to the analyzed sign language data, as well as user emotion information, from a database, and means for transmitting the obtained background information, meaning, and emotion information to the terminal. As a result, when a user who uses sign language uses the electronic payment system, not only is the background information and meaning of the sign language displayed in an easy-to-understand manner, but appropriate feedback according to the user's emotional state is also provided.
[0796] A "user" refers to an individual who operates the system and inputs sign language and emotional data.
[0797] "Sign language data" refers to information that digitizes actions entered by users using sign language.
[0798] "Emotional data" refers to digitized information that represents the emotional state of a user, analyzed from their facial expressions and voice.
[0799] A "terminal" refers to a device used by users to input sign language and emotional data and send that information to a server.
[0800] A "server" refers to a device that analyzes sign language data and emotional data, retrieves the results from a database, and transmits them to a terminal.
[0801] An "emotion recognition algorithm" refers to a computational method for analyzing a user's emotional state from their facial expressions and voice.
[0802] An "image recognition algorithm" refers to a computational method for analyzing sign language data and identifying sign language movements.
[0803] An "identifier" refers to a specific feature or attribute associated with the analyzed sign language data and sentiment data.
[0804] "Background information" refers to information related to sign language, such as its place of origin, history, and cultural background.
[0805] "Meaning" refers to the concept or intention represented by a particular sign language action.
[0806] An "electronic payment system" refers to a system that allows users to conduct online transactions using digital currencies, credits, etc.
[0807] This invention provides a system that enables users who use sign language to smoothly utilize electronic payment systems. This system is implemented with the following configuration and means.
[0808] First, users input sign language using a device such as a smartphone. The device has a built-in camera, which can capture the user's sign language movements as video. Users can also input voice, which allows for the acquisition of facial expressions and vocal information. This data is collected as "sign language data" and "emotional data."
[0809] Next, the terminal sends the collected sign language data and emotion data to the server. An internet connection is used for transmission. The server implements image recognition algorithms and emotion recognition algorithms for analyzing the sign language and emotion data. Specifically, software such as TensorFlow and OpenCV are used. This analyzes the sign language movements, facial expressions, and voice to identify the user's intended meaning.
[0810] When the server analyzes sign language data and emotion data, it retrieves background information and meaning from a database based on the content. This database stores information such as the region of origin, history, and cultural background of the sign language, as well as the meaning of the sign language. The emotion data also includes the user's emotional state, which is also analyzed by the server. Based on the results of this analysis, the server retrieves background information and meaning of the sign language, as well as the user's emotional state, from the database along with a specific identifier.
[0811] Next, the server sends the acquired background information, meaning, and emotional information to the terminal. The terminal displays the received information to the user in an easy-to-understand format. The display format includes text, images, videos, etc., and is provided in a way that is easy for the user to understand. This ensures that when a user of sign language uses the electronic payment system, not only is the background information and meaning of the sign language clearly displayed, but appropriate feedback corresponding to the emotions the user is feeling is also provided at the same time.
[0812] As a concrete example, consider a case where a user performs the sign language for "thank you" in front of their smartphone camera. The device captures the video and sends it to the server. The server analyzes the sign language data to identify it as the sign language for "thank you," and uses an emotion recognition algorithm to detect that the user's emotional state is joy. The server retrieves background information and meaning of the "thank you" sign language, as well as the user's emotional state, from a database and sends it to the device. The device then displays a message to the user stating something like, "'Thank you' is a sign language expression used to express gratitude, and the user's emotional state is 'joy'."
[0813] Examples of prompt messages are as follows:
[0814] example:
[0815] Please generate a Python program that analyzes the following video and recognizes the meaning of the sign language and the user's emotions. The sign language should be returned as Unicode, and the emotions as text.
[0816] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0817] Step 1:
[0818] Users input sign language using their smartphones. The device's camera is used as the input device, capturing the user's sign language movements in real time. The user's facial expressions and voice are also recorded simultaneously. This data is collected as "sign language data" and "emotional data."
[0819] Input: User's sign language gestures, facial expressions, and voice.
[0820] Output: Sign language data, emotion data (captured video and audio)
[0821] Step 2:
[0822] The terminal collects sign language data and emotional data and sends it to the server. An internet connection is used for transmission. Furthermore, the data undergoes necessary preprocessing (e.g., frame resizing and normalization) before being sent to the server.
[0823] Input: Sign language data, emotion data
[0824] Output: Digital data sent to the server
[0825] Step 3:
[0826] The server analyzes sign language data and emotion data. Sign language data is analyzed using an image recognition algorithm (e.g., a TensorFlow model) to identify the gestures and meanings of the signs. Emotion data is analyzed using an emotion recognition algorithm (e.g., a TensorFlow / Keras model) to identify the user's emotional state.
[0827] Input: Sign language data, emotion data
[0828] Output: Analysis results (sign language identification information, emotional state)
[0829] Step 4:
[0830] Based on the analysis results, the server retrieves background information and meanings of sign language, as well as user emotional information, from the database. The database stores the region of origin, history, cultural background, and meaning of each sign language. Appropriate feedback information tailored to the user's emotional state is also retrieved.
[0831] Input: Analysis results (sign language identification information, emotional state)
[0832] Output: Sign language background information, semantic information, and emotional information
[0833] Step 5:
[0834] The server transmits background information, semantic information, and emotional information acquired by the server to the terminal. The data is structured and presented in a user-friendly format. An internet connection is used for transmission.
[0835] Input: Background information, meaning, sentiment information
[0836] Output: Structured data sent to the terminal
[0837] Step 6:
[0838] The device displays received information to the user. The display format includes text, images, and videos, and is provided in a way that allows the user to easily understand the background information, meaning, and emotional state of the sign language.
[0839] Input: Received structured data
[0840] Output: Information displayed to the user (background information, meaning, emotional state)
[0841] Through these steps, users can smoothly utilize the electronic payment system using sign language, and in the process, receive appropriate feedback that reflects the background information and meaning of the sign language, as well as their own emotional state.
[0842] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0843] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0844] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0845] [Third Embodiment]
[0846] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0847] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0848] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0849] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0850] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0851] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0852] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0853] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0854] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0855] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0856] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0857] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0858] This invention relates to a system for learning sign language in a more accessible way, and in particular, a system that helps users input sign language and understand its background information and meaning. The system consists of three main elements: the user, the terminal, and the server.
[0859] Program Processing Overview
[0860] 1. The user enters the sign language.
[0861] Users perform sign language using their device's camera. Alternatively, they can upload pre-recorded sign language video files to their device.
[0862] 2. The device sends sign language data to the server.
[0863] The device saves the sign language video captured by its camera as digital data. Next, it transmits this sign language data to a server via the internet.
[0864] 3. The server analyzes the sign language data.
[0865] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Specifically, it extracts characteristic parts of the sign language and compares them with an existing sign language dictionary.
[0866] 4. The server retrieves background information and meaning.
[0867] The server queries the database using the identified sign language identifier to retrieve background information and meanings for that sign language. For example, it may retrieve information about the origins of the sign language, as well as its cultural and historical background.
[0868] 5. The server sends information to the terminal.
[0869] The server converts the acquired background information and its meaning into structured data and sends it to the terminal.
[0870] 6. The device displays the information.
[0871] The terminal receives background information and meaning from the server and displays it to the user. Display formats include text, images, and videos.
[0872] Specific example
[0873] Example 1: Inputting the sign language for "thank you"
[0874] The user performs the sign language for "thank you" while pointing it at the device's camera. The device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as the sign language for "thank you." The server retrieves background information and meaning of the sign language for "thank you" from its database and sends it to the device. The device then displays this information to the user. Specifically, it might display something like, "The sign language for 'thank you' is used to express gratitude in Japanese."
[0875] Example 2: Inputting regional sign language
[0876] The user inputs sign language specific to the Kansai region into the terminal. The terminal captures this sign language and sends it to the server. The server analyzes the sign language data and recognizes that this sign language is specific to the Kansai region. The server retrieves the regional background and cultural meaning of this sign language from its database and sends it to the terminal. The terminal displays this information to the user. Specifically, it displays, "This sign language represents a greeting unique to people in the Kansai region."
[0877] In this way, by providing users with background information and meanings of sign language data, this system can improve the efficiency of sign language learning and enhance users' familiarity with sign language.
[0878] The following describes the processing flow.
[0879] Step 1: The user enters sign language.
[0880] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[0881] Step 2: The device captures and saves the sign language data.
[0882] The device saves sign language videos captured by the camera as digital data. Similarly, uploaded video files are also saved as digital data.
[0883] Step 3: The device sends the sign language data to the server.
[0884] The terminal transmits the stored sign language data to the server via the internet. During this process, appropriate pre-processing, such as formatting and compression, is performed on the data.
[0885] Step 4: The server receives the sign language data.
[0886] The server receives sign language data sent from the terminal.
[0887] Step 5: The server analyzes the sign language data.
[0888] The server applies an image recognition algorithm to analyze the received sign language data. It extracts characteristic parts of sign language movements and evaluates their similarity by comparing them with an existing sign language dictionary.
[0889] Step 6: The server identifies the sign language identifier.
[0890] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database.
[0891] Step 7: The server retrieves background information and meaning.
[0892] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[0893] Step 8: The server converts background information and meaning into structured data.
[0894] The server converts the acquired background information and meaning into structured data in a format that is easy for the user to read.
[0895] Step 9: The server sends the structured data to the terminal.
[0896] The server sends structured data containing background information and semantics to the terminal.
[0897] Step 10: The device receives and displays information.
[0898] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, videos, etc., and is provided in a way that makes it easy for the user to understand the background information and meaning of sign language.
[0899] In this way, users can understand detailed background information and meaning related to the sign language they input. This process is designed to make learning sign language efficient and accessible.
[0900] (Example 1)
[0901] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0902] When learning sign language, many users face the challenge of understanding the precise meaning and background information of sign language movements. Furthermore, existing sign language learning systems often suffer from low accuracy in analyzing sign language or insufficient acquisition of background information, which hinders the improvement of sign language learning efficiency.
[0903] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0904] In this invention, the server includes means for analyzing transmitted sign language data using a machine learning algorithm, means for obtaining background information and meaning related to the analyzed sign language data from a data store, and means for sending the obtained background information and meaning to a terminal. As a result, when a user inputs sign language, they can quickly obtain the accurate meaning and background information of that sign language.
[0905] A "user" is a person who uses this system to input sign language and receives the analysis results and background information.
[0906] An "input device" is a device used by a user to input sign language, and specifically refers to a camera or touchscreen.
[0907] A "communication device" refers to a device used to transmit input sign language data to a server, and specifically includes wireless communication modules and internet connectivity functions.
[0908] A "server" is a computing device responsible for analyzing sign language data, acquiring background information and meaning, and transmitting this information to users.
[0909] A "machine learning algorithm" is a mathematical method used to analyze sign language data and identify its meaning, specifically referring to Convolutional Neural Networks (CNNs), among others.
[0910] A "data store" refers to a database or storage system used to store background information and meanings of sign language.
[0911] A "display device" is a device used to visually display received information to a user, and specifically refers to displays and monitors.
[0912] A "unique identifier" is a code or number used to uniquely identify a particular sign language, and is used in database queries.
[0913] Modes for carrying out the invention
[0914] This invention relates to a system that helps users input sign language and understand its background information and meaning. The following describes specific embodiments for carrying out the invention.
[0915] Hardware and software to be used
[0916] 1. User
[0917] The user's role is to input sign language. This is done using input devices equipped with cameras and touchscreens.
[0918] 2. Terminal
[0919] The terminal's role is to capture sign language data entered by the user and send it to the server. The terminal's communication equipment includes a wireless communication module and internet connectivity.
[0920] The terminal includes video capture devices (e.g., cameras) and data display devices (e.g., displays).
[0921] 3. Server
[0922] The server's role is to analyze sign language data, extract background information and meaning, and provide it to the user. Sign language analysis utilizes image recognition algorithms such as Convolutional Neural Networks (CNNs). Specifically, machine learning frameworks like TensorFlow and PyTorch are used.
[0923] The server also has a data transmission function to send the acquired background information and meaning to the terminal.
[0924] 4. Datastore
[0925] Data stores are used to store background information and meanings of sign language. SQL and NoSQL databases are typical examples.
[0926] Specific example
[0927] Example 1: Inputting "thank you" in sign language
[0928] 1. The user performs the sign language for "thank you" while pointing it at the device's camera.
[0929] The user performs the sign language for "thank you" in front of the camera.
[0930] 2. The device captures this video and sends it to the server as sign language data.
[0931] The device's camera captures the video frame by frame and saves it as digital data. The data is then sent to a server.
[0932] 3. The server analyzes the sign language data and identifies it as the sign for "thank you."
[0933] The server analyzes the received data using an image recognition algorithm and identifies it as the sign language for "thank you."
[0934] 4. The server retrieves background information and meaning of the sign language for "thank you" from the database and displays the retrieved information on the terminal.
[0935] The server retrieves information related to the sign language expression for "thank you" from its database. For example, this information may include details about how gratitude is expressed in Japanese culture. The retrieved information is then sent to the terminal as structured data (e.g., in JSON format).
[0936] 5. Display the information received by the device to the user.
[0937] The device visually displays the information it receives to the user. For example, it might display information such as, "The sign language for 'thank you' is used to express gratitude in Japanese."
[0938] Example of a prompt
[0939] The following is an example of a prompt statement for inputting a system description into the generating AI model.
[0940] Prompt message:
[0941] I would like to describe a system for learning sign language. This system helps users input sign language and understand its meaning and background information. Please explain the program's process in natural language, following the steps below.
[0942] procedure:
[0943] 1. The user inputs the sign language. The user either performs the sign language using their device's camera or uploads a pre-recorded video file of the sign language.
[0944] 2. The device sends sign language data to the server. The device saves the sign language video captured by the camera as digital data and sends it to the server via the internet.
[0945] 3. The server analyzes the sign language data. The server uses an image recognition algorithm to analyze what the sign language data means and compares it to an existing sign language dictionary.
[0946] 4. The server retrieves background information and meaning. The server queries the database using the identified sign language identifier to retrieve background information and meaning for that sign language.
[0947] 5. The server sends the information to the terminal. The server converts the acquired background information and meaning into structured data and sends it to the terminal.
[0948] 6. The terminal displays information. The terminal receives background information and meaning sent from the server and displays it to the user. Display formats include text, images, videos, etc.
[0949] Specific example:
[0950] Please explain the processing steps using the sign language for "thank you" as an example.
[0951] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0952] Specific flow of program processing
[0953] Processing steps
[0954] Step 1:
[0955] The user inputs sign language. The user either performs sign language using the camera on their device or uploads a pre-recorded video file of sign language. The input data is a video of sign language. This video of sign language becomes the raw data used for analysis in subsequent processing.
[0956] Step 2:
[0957] The device captures sign language data and saves it as digital data. The device's camera captures video frame by frame and saves this as digital data (e.g., MP4 format) to internal storage. The captured data is used for subsequent analysis.
[0958] Step 3:
[0959] The terminal transmits the stored sign language data to the server via a communication device. The terminal uses an internet connection to transmit the stored sign language data to the server. The input data is the captured sign language video data, and the output is the data transmitted to the server.
[0960] Step 4:
[0961] The server receives sign language data and analyzes it using a machine learning algorithm. The server then analyzes the received data using an image recognition algorithm (e.g., a Convolutional Neural Network). This analysis extracts the features of the sign language and converts them into an identifiable format. The input data is video data of sign language, and the output is the analyzed sign language data.
[0962] Step 5:
[0963] The server identifies a unique identifier based on the analyzed sign language data and retrieves background information and meaning from the data store. The server queries the database based on the analysis results to retrieve background information and meaning for the corresponding sign language. The input data is the analyzed sign language data, and the output is data related to background information and meaning.
[0964] Step 6:
[0965] The server converts the acquired information into structured data (e.g., JSON format) and sends it to the terminal. The server uses a data transmission function to send the acquired information to the terminal. The input data is background information and semantic data, and the output is the structured data sent to the terminal.
[0966] Step 7:
[0967] The terminal displays the information it receives to the user. The terminal's display device analyzes the received structured data and displays it visually to the user. Display formats include text, images, and videos. The input data is structured data, and the output is the information displayed to the user.
[0968] Detailed explanation of operation
[0969] Step 1:
[0970] Users perform the sign language for "thank you" using their camera. Users can also upload the recorded video file.
[0971] Step 2:
[0972] The device's camera captures video of the user performing sign language. This video is saved as digital data to the internal storage.
[0973] Step 3:
[0974] The device transmits the stored sign language video data to the server via the internet.
[0975] Step 4:
[0976] The server analyzes the received sign language video data using a Convolutional Neural Network. It extracts characteristic patterns from the sign language and converts them into specific identifiers.
[0977] Step 5:
[0978] The server uses the identified identifier to query the database and retrieve background information and meanings for the corresponding sign language.
[0979] Step 6:
[0980] The server converts the background information and meaning retrieved from the database into structured data in JSON format and sends it to the terminal.
[0981] Step 7:
[0982] The structured data received by the terminal is displayed on the display device, and information such as "The sign language for 'thank you' is used to express gratitude" is presented to the user.
[0983] (Application Example 1)
[0984] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0985] Conventional sign language learning systems not only teach sign language but also lack systems that enable its practical application in real-world work environments. In particular, it is difficult for deaf individuals to give work instructions to machines and robots using sign language in factories and manufacturing sites. In such environments, a work instruction system that utilizes sign language is needed.
[0986] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0987] In this invention, the server includes means for analyzing sign language data, means for obtaining background information and meaning related to the analyzed sign language data from a database, and means for identifying work instructions based on the obtained background information and meaning and transmitting them to a device. This enables reliable transmission of work instructions using sign language.
[0988] A "user" refers to a person who uses the system to input sign language and issue work instructions.
[0989] "Sign language" refers to a system of gestures and hand movements used by people with hearing impairments to communicate.
[0990] "Means" refers to the methods, devices, and components used to realize each function of a system.
[0991] "Data" refers to information including video recordings of sign language and the results of their analysis that the system processes.
[0992] "Server" refers to a central processing unit that performs sign language data analysis and generates work instructions.
[0993] "Analysis" refers to the processing of an algorithm that receives sign language data, understands its content, and identifies its meaning.
[0994] "Background information" refers to the cultural and historical contextual information associated with the analyzed sign language.
[0995] "Meaning" refers to the content or intention that sign language is meant to convey.
[0996] A "database" refers to a collection of information that stores sign language, its background information, and its meaning in an associated manner.
[0997] A "terminal" refers to a device used by a user to input sign language and receive and display the analysis results.
[0998] "Equipment" refers to a device that receives and executes work instructions generated from analyzed sign language data.
[0999] "Work instructions" refer to specific operational instructions that robots and other automated equipment should perform based on the analyzed sign language.
[1000] "Execution" refers to the physical actions or processes that a device performs based on the work instructions it receives.
[1001] Modes for carrying out the invention
[1002] The system that realizes this invention involves a user inputting sign language, analyzing that data to generate relevant work instructions, and transmitting those instructions to a device for execution. This system is mainly composed of the following hardware and software.
[1003] hardware
[1004] Camera module: This utilizes cameras attached to factory robots, as well as the built-in cameras of smartphones and tablets. This allows for the capture of user sign language.
[1005] Factory robots: Specifically, this includes automation equipment from companies such as KUKA, ABB, and Fanuc. These robots are responsible for receiving and executing work instructions.
[1006] software
[1007] Image recognition algorithm: This algorithm analyzes sign language data using tools such as TensorFlow and OpenCV. It extracts features from sign language and identifies what those signs mean.
[1008] Sign language database: MongoDB or PostgreSQL is used. This database stores the meaning and background information of sign language.
[1009] Server-side program: Uses Node.js or Python (Flask) to analyze sign language data and generate work instructions.
[1010] Frontend: JavaScript (React.js) is used to implement the user interface.
[1011] Specific examples of actions
[1012] 1. User inputs sign language: Users perform sign language towards a camera attached to the factory robot. Alternatively, they can input sign language from a smartphone or tablet. The sign language performed by the user mainly corresponds to work instructions such as "assemble" and "move".
[1013] Example prompt: "Perform the following sign language towards the camera and enter a work instruction: 'Move,' the robot will move."
[1014] 2. Data Transmission: The collected sign language data is sent to the server using WebSocket. The server divides the received sign language data into frames and prepares them for analysis.
[1015] 3. Data Analysis: An image recognition algorithm using TensorFlow is executed on the server side to analyze the received sign language data. This identifies the meaning of the sign language.
[1016] 4. Database query: The server uses the identified sign language identifier to retrieve the corresponding work instruction from a database such as MongoDB.
[1017] 5. Sending work instructions: The server sends the generated work instructions to the equipment. The equipment operates automatically based on the received work instructions.
[1018] Details of specific examples
[1019] For example, if a user performs the sign language for "move" while pointing it at the camera, the camera module captures the sign language. This data is sent to a server in real time and analyzed by TensorFlow. Based on the analysis results, the task instruction "move" is identified and retrieved from the database. The task instruction is then sent to the robot, which performs the specified movement.
[1020] This will enable people with hearing impairments to efficiently operate equipment in factories and manufacturing sites using sign language.
[1021] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1022] Step 1:
[1023] The user enters sign language.
[1024] The user performs sign language towards a camera attached to a factory robot or towards a smartphone or tablet. The entered sign language is captured as video and stored as digital data for use in the next processing step.
[1025] Input: User's sign language video
[1026] Output: Captured sign language video data
[1027] Step 2:
[1028] The device sends sign language data to the server.
[1029] The terminal (a robot or smartphone with a built-in camera module) divides the captured sign language video into frames in real time and sends them to the server using WebSocket.
[1030] Input: Captured sign language video data
[1031] Output: Sign language video frames sent to the server
[1032] Step 3:
[1033] The server analyzes the sign language data.
[1034] The server analyzes the received sign language data using a TensorFlow-based image recognition algorithm. It extracts each feature point of the sign language and performs analysis to determine what the sign language means.
[1035] Input: Sent sign language video frame
[1036] Output: Meaning of the analyzed sign language (identified sign language identifier)
[1037] Step 4:
[1038] The server retrieves background information and meaning from the database.
[1039] The server uses the identified sign language identifier to query a MongoDB or PostgreSQL database to retrieve background information, meaning, and corresponding work instructions for that sign language.
[1040] Input: Identified sign language identifier
[1041] Output: Background information, meaning, work instructions
[1042] Step 5:
[1043] The server sends work instructions to the equipment.
[1044] The server converts the acquired work instructions into structured data in JSON format and transmits it to equipment such as factory robots via the internet.
[1045] Input: Work instructions
[1046] Output: Work instructions sent to the device
[1047] Step 6:
[1048] The equipment performs tasks based on the work instructions it receives.
[1049] Factory automation equipment and robots analyze received work instructions and perform actions based on those instructions. For example, if the instruction is to move, the robot will move to the specified location.
[1050] Input: Received work instructions
[1051] Output: Performed actions (e.g., robot movement)
[1052] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1053] This invention relates to a system for learning sign language in a more accessible way, and in particular, a system that helps users input sign language and understand its background information, meaning, and even their emotions. The system consists of four main elements: the user, the terminal, the server, and the emotion engine.
[1054] Program Processing Overview
[1055] 1. The user enters the sign language.
[1056] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[1057] In this process, the device uses an emotion engine to acquire emotional data from the user's facial expressions and voice.
[1058] 2. The device sends sign language data and emotion data to the server.
[1059] The device stores sign language video captured by the camera and emotional data analyzed by the emotion engine as digital data.
[1060] Next, this sign language data and emotion data are sent to a server via the internet.
[1061] 3. The server analyzes the sign language data.
[1062] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Specifically, it extracts characteristic parts of the sign language and evaluates its similarity by comparing it with an existing sign language dictionary.
[1063] 4. The server identifies the sign language identifier.
[1064] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database.
[1065] 5. The server retrieves background information and meaning.
[1066] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[1067] 6. The server analyzes the emotional data.
[1068] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the user's emotional state is joyful when they input the sign language for "thank you."
[1069] 7. The server converts background information, semantic, and sentiment data into structured data.
[1070] The server converts the sign language background information, meaning, and emotional data into structured data and arranges it in a user-friendly format.
[1071] 8. The server sends the structured data to the terminal.
[1072] The server sends structured data regarding background information, meaning, and sentiment to the terminal.
[1073] 9. The device receives and displays information.
[1074] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, and videos, and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language.
[1075] Specific example
[1076] Example 1: Input of the sign language "thank you" and analysis of emotional data
[1077] The user performs the sign language for "thank you" while pointing it at the device's camera. The device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as the sign language for "thank you." It also uses an emotion engine to detect that the user's emotional state is joy. The server retrieves background information and meaning of the sign language for "thank you" from its database, adds the emotional state, and sends it to the device. The device then displays a message to the user stating something like, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1078] Example 2: Input of region-specific sign language and analysis of emotional data
[1079] The user inputs sign language specific to the Kansai region into the terminal. The terminal captures this sign language and sends it to the server along with emotion data. The server analyzes the sign language data and emotion data and recognizes that the sign language is specific to the Kansai region and that the user's emotional state is surprise or confusion. The server sends the regional background and cultural meaning of the sign language, as well as the emotional state, to the terminal. The terminal displays, "This sign language represents a greeting specific to the Kansai region, and the user's emotional state is 'surprise' or 'confusion'."
[1080] In this way, this system can improve the efficiency of sign language learning and the user's sense of familiarity by providing integrated background information, meaning, and the user's emotional state in sign language data.
[1081] The following describes the processing flow.
[1082] Step 1: The user enters sign language.
[1083] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[1084] The device captures sign language movements in real time using its camera.
[1085] Step 2: The device captures and saves sign language data and emotion data.
[1086] The device utilizes an emotion engine to analyze the user's emotional state in real time using facial recognition and voice analysis technologies.
[1087] These analysis results will be saved as digital data and managed together with the sign language data.
[1088] Step 3: The device sends sign language data and emotion data to the server.
[1089] The terminal transmits the stored sign language data and emotional data to the server via the internet. During this process, appropriate pre-processing, such as formatting and compression, is performed on the data.
[1090] Step 4: The server receives the sign language data.
[1091] The server receives sign language data and emotion data sent from the terminal.
[1092] Step 5: The server analyzes the sign language data.
[1093] The server analyzes the received sign language data and uses image recognition algorithms to identify what the sign language actions mean. For example, it extracts characteristic parts of the sign language actions and compares them with existing sign language dictionaries.
[1094] Step 6: The server identifies the sign language identifier.
[1095] The server identifies the corresponding sign language identifier based on the results of the image recognition algorithm. This identifier corresponds to a sign language entry in the database.
[1096] Step 7: The server retrieves background information and meaning.
[1097] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[1098] Step 8: The server analyzes the emotion data.
[1099] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it determines whether the user is exhibiting emotions such as "joy" or "surprise."
[1100] Step 9: The server structures all the data.
[1101] The server integrates background information, meaning, and emotional data from sign language and structures it into a format that is easy for users to understand.
[1102] Step 10: The server sends the structured data to the terminal.
[1103] The server sends integrated structured data to the terminal.
[1104] Step 11: The device receives and displays information.
[1105] The terminal receives structured data sent from the server and displays it to the user. The displayed content includes background information and meaning of sign language, as well as the user's emotional state. For example, it may be displayed visually in an easy-to-understand way using text, images, and videos.
[1106] Specific example
[1107] Example 1: Input of the sign language "thank you" and analysis of emotional data
[1108] 1. The user performs the sign language for "thank you" while pointing it at the device's camera.
[1109] 2. The device captures video and saves sign language data and emotion data analyzed by the emotion engine. The emotion engine detects the emotion of "joy" from the user's facial expressions and voice.
[1110] 3. The device sends sign language data and emotion data to the server.
[1111] 4. The server receives the sign language data.
[1112] 5. The server analyzes the sign language data using an image recognition algorithm.
[1113] 6. The server identifies the sign language identifier.
[1114] 7. The server retrieves background information and meaning of the sign language for "thank you" from the database.
[1115] 8. The server analyzes the emotion data and confirms that the user is experiencing the emotion of "joy."
[1116] 9. The server formats all data into structured data and sends it to the terminal.
[1117] 10. The device receives the data and displays the message, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1118] Example 2: Input of region-specific sign language and analysis of emotional data
[1119] 1. The user inputs sign language specific to the Kansai region into the device.
[1120] 2. The device captures sign language and sends it to the server along with emotion data. The emotion engine detects the emotion of "surprise" from the user's facial expressions.
[1121] 3. The server analyzes the sign language data and emotion data to confirm that it is a regionally specific sign language and to verify the emotional state.
[1122] 4. The server retrieves regional background and cultural significance from the database.
[1123] 5. The server integrates this data and sends it to the terminal.
[1124] 6. The device displays the message, "This sign language represents a greeting unique to the Kansai region, and the user's emotional state is 'surprise'."
[1125] Through this system, users can understand not only the background information and meaning of sign language, but also their own emotional state, making sign language learning more accessible.
[1126] (Example 2)
[1127] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1128] Conventional sign language learning systems have the drawback of not being able to simultaneously understand not only the background information and meaning of sign language, but also the emotional state of the user performing the sign language. Furthermore, there is a need to make learning sign language more accessible and effective by reflecting the user's emotional state when inputting sign language, but an effective system for this purpose has not existed.
[1129] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing transmitted sign language data and emotion data, means for obtaining background information and meaning related to the analyzed sign language data from a database, and means for evaluating the user's emotional state from the analyzed emotion data. This makes it possible to simultaneously grasp not only the background information and meaning of the sign language, but also the user's emotional state when performing the sign language.
[1130] A "user" refers to a person who uses the system to input sign language.
[1131] "Means of inputting sign language" refers to the interface (such as a camera or touchscreen) that a user uses to input sign language into a device.
[1132] "Sign language data" refers to information that digitally records the sign language actions entered by the user.
[1133] "Emotional data" refers to digital information that indicates the emotional state of a user, analyzed from their facial expressions, voice, and other data.
[1134] A "server" refers to a central computing device that performs data analysis, processing, and storage.
[1135] "Means of transmission" refers to the means of communication used to transfer data from a terminal to a server.
[1136] "Means of analysis" refers to algorithms and technologies (for example, image recognition algorithms and sentiment analysis technologies) that the server uses to process sign language data and sentiment data and analyze their content.
[1137] "Background information" refers to historical, cultural, and regional information related to sign language.
[1138] "Meaning" refers to the concept or intention expressed by the movements in sign language.
[1139] A "database" refers to an information management system that stores sign language data, background information, semantic data, and emotional data, making them accessible as needed.
[1140] A "specific identifier" refers to a code or label used to uniquely identify the type or meaning of sign language.
[1141] "Structured data" refers to data that follows a specific format and includes formalized data such as analyzed sign language information and emotional states.
[1142] A "terminal" refers to a device (such as a smartphone or tablet) used by a user to input sign language, receive data, and display it.
[1143] "Means of display" refers to the means by which a terminal visually presents analysis results and acquired information to the user.
[1144] This invention relates to a system for learning sign language in a more accessible way. The system helps users input sign language and understand its background information, meaning, and even their emotions. The system consists of four main elements: the user, the terminal, the server, and the emotion engine.
[1145] System Overview
[1146] 1. The user enters the sign language.
[1147] The user performs sign language actions in front of the device's camera. For example, to input the sign language for "thank you," the user lightly presses their palms in front of their chest. Alternatively, they can upload a pre-recorded video file of sign language to the device. In this case, the device's camera captures the user's sign language movements, and an emotion engine is used to obtain emotional data from the user's facial expressions and voice.
[1148] 2. The device sends sign language data and emotion data to the server.
[1149] The device captures sign language video with its camera and saves it as digital data. It also collects emotional data analyzed by an emotion engine. The captured sign language data and emotional data are sent together as a data packet to the server via the internet. Examples of emotion engines that can be used include "OpenCV" and "Microsoft Azure Face API".
[1150] 3. The server analyzes the sign language data.
[1151] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Examples of image recognition algorithms used include "TensorFlow" and "PyTorch". It extracts characteristic parts of the sign language and evaluates similarity by comparing them with existing sign language dictionaries.
[1152] 4. The server identifies the sign language identifier.
[1153] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database. For example, the sign language for "thank you" is identified as identifier "S001".
[1154] 5. The server retrieves background information and meaning.
[1155] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[1156] 6. The server analyzes the emotional data.
[1157] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the emotional state of a user when they input the sign language for "thank you" is "joy." Possible emotional analysis platforms used include "IBM Watson" and "Amazon Rekognition."
[1158] 7. The server converts background information, semantic, and sentiment data into structured data.
[1159] The server converts the sign language background information, meaning, and emotional data into structured data and arranges it in a user-friendly format. For example, it converts it into a structured data format such as JSON.
[1160] 8. The server sends the structured data to the terminal.
[1161] The server sends structured data containing background information, semantic, and sentiment data to the terminal. Communication protocols used include HTTP and WebSocket.
[1162] 9. The device receives and displays information.
[1163] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, and videos, and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language. For example, the terminal screen may display the message, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1164] Specific example
[1165] Example 1: Input of the sign language "thank you" and analysis of emotional data
[1166] The user performs the sign language for "thank you" towards the device's camera, and the device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as identifier "S001". It also uses an emotion engine to detect that the user's emotional state is "joy". The server retrieves background information and meaning of the "thank you" sign language from the database, adds the emotional state, and sends it to the device. The device displays to the user, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1167] Example 2: Input of region-specific sign language and analysis of emotional data
[1168] The user inputs a sign language specific to the Kansai region into the terminal, and the terminal captures this sign language and sends it to the server along with emotion data. The server analyzes the sign language data and emotion data and recognizes that the sign language is specific to the Kansai region and that the user's emotional state is "surprise" or "confusion". The server sends the regional background and cultural meaning of the sign language, as well as the emotional state, to the terminal. The terminal displays, "This sign language expresses a greeting specific to the Kansai region, and the user's emotional state is 'surprise' or 'confusion'."
[1169] Example of a prompt
[1170] 1. "Please analyze the emotions expressed when performing sign language and provide this information along with the background context."
[1171] 2. What system can analyze the sign language for "thank you" and display the emotional state?
[1172] In this way, this system can improve the efficiency of sign language learning and the user's sense of familiarity by providing integrated background information, meaning, and the user's emotional state in sign language data.
[1173] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1174] Step 1:
[1175] The user enters sign language.
[1176] The user performs sign language actions in front of the device's camera. For example, to sign "thank you," the user lightly presses their palms in front of their chest. Alternatively, they can upload a pre-recorded sign language video file to the device. Input can be real-time sign language video or uploaded video files. Output is the device acquiring this sign language video data.
[1177] Step 2:
[1178] The device sends sign language data and emotion data to the server.
[1179] The device stores sign language video captured by its camera as digital data and uses an emotion engine to analyze emotional data from the user's facial expressions and voice. For example, the emotion engine uses facial recognition technology to analyze the degree of the user's smile and the tone of their voice. The input consists of the captured sign language video data and the results of the emotion analysis. As output, this data is sent to the server as data packets.
[1180] Step 3:
[1181] The server analyzes the sign language data.
[1182] The server inputs the received sign language data into an image recognition algorithm, extracts the characteristic parts of the sign language, and begins analysis. Examples of algorithms used include "TensorFlow" and "PyTorch." The input is sign language video data. The output is a sign language feature vector.
[1183] Step 4:
[1184] The server identifies the sign language identifier.
[1185] The server compares the generated feature vectors with existing sign language dictionaries to find the most similar sign language data. Each identified sign language data is assigned a unique identifier. The input consists of feature vectors and a sign language dictionary. The output is a sign language identifier (e.g., "S001").
[1186] Step 5:
[1187] The server retrieves background information and meaning.
[1188] The server uses the identified sign language identifier to retrieve background information (such as the region of origin, history, and cultural background) and its meaning from the database. The input is the sign language identifier. The output is the background information and meaning of the sign language.
[1189] Step 6:
[1190] The server analyzes the emotional data.
[1191] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the user's emotional state when they sign "thank you" is "joy." The input is emotional data, and the output is an evaluation of the user's emotional state.
[1192] Step 7:
[1193] The server converts background information, semantic, and sentiment data into structured data.
[1194] The server converts the sign language background information, meaning, and analyzed sentiment data into structured data (e.g., JSON format) and arranges it in a user-friendly format. Input includes background information, meaning, and sentiment data. The output is structured data.
[1195] Step 8:
[1196] The server sends structured data to the terminal.
[1197] The server sends structured data to the terminal. Communication protocols used include "HTTP" and "WebSocket". The input is structured data. The output is this data sent to the terminal.
[1198] Step 9:
[1199] The device receives and displays information.
[1200] The terminal receives structured data sent from the server and passes it to the UI for display to the user. The display format includes text, images, videos, etc., and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language. The input is structured data. The output is information that is displayed to the user. For example, the terminal screen might display "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1201] (Application Example 2)
[1202] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1203] This invention relates to a system for improving the ease of use and usability of electronic payment systems for users who use sign language. Conventional electronic payment systems are based on voice or text-based input, making them difficult to operate for users who use sign language. Furthermore, they lacked the function to analyze the user's emotional state and provide appropriate feedback accordingly. Therefore, there is a need to provide a system that integrates sign language input and emotion analysis to enable users who use sign language to comfortably make electronic payments.
[1204] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing input sign language data and user emotion data, means for obtaining background information and meaning related to the analyzed sign language data, as well as user emotion information, from a database, and means for transmitting the obtained background information, meaning, and emotion information to the terminal. As a result, when a user who uses sign language uses the electronic payment system, not only is the background information and meaning of the sign language displayed in an easy-to-understand manner, but appropriate feedback according to the user's emotional state is also provided.
[1205] A "user" refers to an individual who operates the system and inputs sign language and emotional data.
[1206] "Sign language data" refers to information that digitizes actions entered by users using sign language.
[1207] "Emotional data" refers to digitized information that represents the emotional state of a user, analyzed from their facial expressions and voice.
[1208] A "terminal" refers to a device used by users to input sign language and emotional data and send that information to a server.
[1209] A "server" refers to a device that analyzes sign language data and emotional data, retrieves the results from a database, and transmits them to a terminal.
[1210] An "emotion recognition algorithm" refers to a computational method for analyzing a user's emotional state from their facial expressions and voice.
[1211] An "image recognition algorithm" refers to a computational method for analyzing sign language data and identifying sign language movements.
[1212] An "identifier" refers to a specific feature or attribute associated with the analyzed sign language data and sentiment data.
[1213] "Background information" refers to information related to sign language, such as its place of origin, history, and cultural background.
[1214] "Meaning" refers to the concept or intention represented by a particular sign language action.
[1215] An "electronic payment system" refers to a system that allows users to conduct online transactions using digital currencies, credits, etc.
[1216] This invention provides a system that enables users who use sign language to smoothly utilize electronic payment systems. This system is implemented with the following configuration and means.
[1217] First, users input sign language using a device such as a smartphone. The device has a built-in camera, which can capture the user's sign language movements as video. Users can also input voice, which allows for the acquisition of facial expressions and vocal information. This data is collected as "sign language data" and "emotional data."
[1218] Next, the terminal sends the collected sign language data and emotion data to the server. An internet connection is used for transmission. The server implements image recognition algorithms and emotion recognition algorithms for analyzing the sign language and emotion data. Specifically, software such as TensorFlow and OpenCV are used. This analyzes the sign language movements, facial expressions, and voice to identify the user's intended meaning.
[1219] When the server analyzes sign language data and emotion data, it retrieves background information and meaning from a database based on the content. This database stores information such as the region of origin, history, and cultural background of the sign language, as well as the meaning of the sign language. The emotion data also includes the user's emotional state, which is also analyzed by the server. Based on the results of this analysis, the server retrieves background information and meaning of the sign language, as well as the user's emotional state, from the database along with a specific identifier.
[1220] Next, the server sends the acquired background information, meaning, and emotional information to the terminal. The terminal displays the received information to the user in an easy-to-understand format. The display format includes text, images, videos, etc., and is provided in a way that is easy for the user to understand. This ensures that when a user of sign language uses the electronic payment system, not only is the background information and meaning of the sign language clearly displayed, but appropriate feedback corresponding to the emotions the user is feeling is also provided at the same time.
[1221] As a concrete example, consider a case where a user performs the sign language for "thank you" in front of their smartphone camera. The device captures the video and sends it to the server. The server analyzes the sign language data to identify it as the sign language for "thank you," and uses an emotion recognition algorithm to detect that the user's emotional state is joy. The server retrieves background information and meaning of the "thank you" sign language, as well as the user's emotional state, from a database and sends it to the device. The device then displays a message to the user stating something like, "'Thank you' is a sign language expression used to express gratitude, and the user's emotional state is 'joy'."
[1222] Examples of prompt messages are as follows:
[1223] example:
[1224] Please generate a Python program that analyzes the following video and recognizes the meaning of the sign language and the user's emotions. The sign language should be returned as Unicode, and the emotions as text.
[1225] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1226] Step 1:
[1227] Users input sign language using their smartphones. The device's camera is used as the input device, capturing the user's sign language movements in real time. The user's facial expressions and voice are also recorded simultaneously. This data is collected as "sign language data" and "emotional data."
[1228] Input: User's sign language gestures, facial expressions, and voice.
[1229] Output: Sign language data, emotion data (captured video and audio)
[1230] Step 2:
[1231] The terminal collects sign language data and emotional data and sends it to the server. An internet connection is used for transmission. Furthermore, the data undergoes necessary preprocessing (e.g., frame resizing and normalization) before being sent to the server.
[1232] Input: Sign language data, emotion data
[1233] Output: Digital data sent to the server
[1234] Step 3:
[1235] The server analyzes sign language data and emotion data. Sign language data is analyzed using an image recognition algorithm (e.g., a TensorFlow model) to identify the gestures and meanings of the signs. Emotion data is analyzed using an emotion recognition algorithm (e.g., a TensorFlow / Keras model) to identify the user's emotional state.
[1236] Input: Sign language data, emotion data
[1237] Output: Analysis results (sign language identification information, emotional state)
[1238] Step 4:
[1239] Based on the analysis results, the server retrieves background information and meanings of sign language, as well as user emotional information, from the database. The database stores the region of origin, history, cultural background, and meaning of each sign language. Appropriate feedback information tailored to the user's emotional state is also retrieved.
[1240] Input: Analysis results (sign language identification information, emotional state)
[1241] Output: Sign language background information, semantic information, and emotional information
[1242] Step 5:
[1243] The server transmits background information, semantic information, and emotional information acquired by the server to the terminal. The data is structured and presented in a user-friendly format. An internet connection is used for transmission.
[1244] Input: Background information, meaning, sentiment information
[1245] Output: Structured data sent to the terminal
[1246] Step 6:
[1247] The device displays received information to the user. The display format includes text, images, and videos, and is provided in a way that allows the user to easily understand the background information, meaning, and emotional state of the sign language.
[1248] Input: Received structured data
[1249] Output: Information displayed to the user (background information, meaning, emotional state)
[1250] Through these steps, users can smoothly utilize the electronic payment system using sign language, and in the process, receive appropriate feedback that reflects the background information and meaning of the sign language, as well as their own emotional state.
[1251] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1252] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1253] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1254] [Fourth Embodiment]
[1255] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1256] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1257] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1258] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1259] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1260] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1261] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1262] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1263] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1264] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1265] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1266] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1267] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1268] This invention relates to a system for learning sign language in a more accessible way, and in particular, a system that helps users input sign language and understand its background information and meaning. The system consists of three main elements: the user, the terminal, and the server.
[1269] Program Processing Overview
[1270] 1. The user enters the sign language.
[1271] Users perform sign language using their device's camera. Alternatively, they can upload pre-recorded sign language video files to their device.
[1272] 2. The device sends sign language data to the server.
[1273] The device saves the sign language video captured by its camera as digital data. Next, it transmits this sign language data to a server via the internet.
[1274] 3. The server analyzes the sign language data.
[1275] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Specifically, it extracts characteristic parts of the sign language and compares them with an existing sign language dictionary.
[1276] 4. The server retrieves background information and meaning.
[1277] The server queries the database using the identified sign language identifier to retrieve background information and meanings for that sign language. For example, it may retrieve information about the origins of the sign language, as well as its cultural and historical background.
[1278] 5. The server sends information to the terminal.
[1279] The server converts the acquired background information and its meaning into structured data and sends it to the terminal.
[1280] 6. The device displays the information.
[1281] The terminal receives background information and meaning from the server and displays it to the user. Display formats include text, images, and videos.
[1282] Specific example
[1283] Example 1: Inputting the sign language for "thank you"
[1284] The user performs the sign language for "thank you" while pointing it at the device's camera. The device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as the sign language for "thank you." The server retrieves background information and meaning of the sign language for "thank you" from its database and sends it to the device. The device then displays this information to the user. Specifically, it might display something like, "The sign language for 'thank you' is used to express gratitude in Japanese."
[1285] Example 2: Inputting regional sign language
[1286] The user inputs sign language specific to the Kansai region into the terminal. The terminal captures this sign language and sends it to the server. The server analyzes the sign language data and recognizes that this sign language is specific to the Kansai region. The server retrieves the regional background and cultural meaning of this sign language from its database and sends it to the terminal. The terminal displays this information to the user. Specifically, it displays, "This sign language represents a greeting unique to people in the Kansai region."
[1287] In this way, by providing users with background information and meanings of sign language data, this system can improve the efficiency of sign language learning and enhance users' familiarity with sign language.
[1288] The following describes the processing flow.
[1289] Step 1: The user enters sign language.
[1290] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[1291] Step 2: The device captures and saves the sign language data.
[1292] The device saves sign language videos captured by the camera as digital data. Similarly, uploaded video files are also saved as digital data.
[1293] Step 3: The device sends the sign language data to the server.
[1294] The terminal transmits the stored sign language data to the server via the internet. During this process, appropriate pre-processing, such as formatting and compression, is performed on the data.
[1295] Step 4: The server receives the sign language data.
[1296] The server receives sign language data sent from the terminal.
[1297] Step 5: The server analyzes the sign language data.
[1298] The server applies an image recognition algorithm to analyze the received sign language data. It extracts characteristic parts of sign language movements and evaluates their similarity by comparing them with an existing sign language dictionary.
[1299] Step 6: The server identifies the sign language identifier.
[1300] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database.
[1301] Step 7: The server retrieves background information and meaning.
[1302] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[1303] Step 8: The server converts background information and meaning into structured data.
[1304] The server converts the acquired background information and meaning into structured data in a format that is easy for the user to read.
[1305] Step 9: The server sends the structured data to the terminal.
[1306] The server sends structured data containing background information and semantics to the terminal.
[1307] Step 10: The device receives and displays information.
[1308] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, videos, etc., and is provided in a way that makes it easy for the user to understand the background information and meaning of sign language.
[1309] In this way, users can understand detailed background information and meaning related to the sign language they input. This process is designed to make learning sign language efficient and accessible.
[1310] (Example 1)
[1311] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1312] When learning sign language, many users face the challenge of understanding the precise meaning and background information of sign language movements. Furthermore, existing sign language learning systems often suffer from low accuracy in analyzing sign language or insufficient acquisition of background information, which hinders the improvement of sign language learning efficiency.
[1313] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1314] In this invention, the server includes means for analyzing transmitted sign language data using a machine learning algorithm, means for obtaining background information and meaning related to the analyzed sign language data from a data store, and means for sending the obtained background information and meaning to a terminal. As a result, when a user inputs sign language, they can quickly obtain the accurate meaning and background information of that sign language.
[1315] A "user" is a person who uses this system to input sign language and receives the analysis results and background information.
[1316] An "input device" is a device used by a user to input sign language, and specifically refers to a camera or touchscreen.
[1317] A "communication device" refers to a device used to transmit input sign language data to a server, and specifically includes wireless communication modules and internet connectivity functions.
[1318] A "server" is a computing device responsible for analyzing sign language data, acquiring background information and meaning, and transmitting this information to users.
[1319] A "machine learning algorithm" is a mathematical method used to analyze sign language data and identify its meaning, specifically referring to Convolutional Neural Networks (CNNs), among others.
[1320] A "data store" refers to a database or storage system used to store background information and meanings of sign language.
[1321] A "display device" is a device used to visually display received information to a user, and specifically refers to displays and monitors.
[1322] A "unique identifier" is a code or number used to uniquely identify a particular sign language, and is used in database queries.
[1323] Modes for carrying out the invention
[1324] This invention relates to a system that helps users input sign language and understand its background information and meaning. The following describes specific embodiments for carrying out the invention.
[1325] Hardware and software to be used
[1326] 1. User
[1327] The user's role is to input sign language. This is done using input devices equipped with cameras and touchscreens.
[1328] 2. Terminal
[1329] The terminal's role is to capture sign language data entered by the user and send it to the server. The terminal's communication equipment includes a wireless communication module and internet connectivity.
[1330] The terminal includes video capture devices (e.g., cameras) and data display devices (e.g., displays).
[1331] 3. Server
[1332] The server's role is to analyze sign language data, extract background information and meaning, and provide it to the user. Sign language analysis utilizes image recognition algorithms such as Convolutional Neural Networks (CNNs). Specifically, machine learning frameworks like TensorFlow and PyTorch are used.
[1333] The server also has a data transmission function to send the acquired background information and meaning to the terminal.
[1334] 4. Datastore
[1335] Data stores are used to store background information and meanings of sign language. SQL and NoSQL databases are typical examples.
[1336] Specific example
[1337] Example 1: Inputting "thank you" in sign language
[1338] 1. The user performs the sign language for "thank you" while pointing it at the device's camera.
[1339] The user performs the sign language for "thank you" in front of the camera.
[1340] 2. The device captures this video and sends it to the server as sign language data.
[1341] The device's camera captures the video frame by frame and saves it as digital data. The data is then sent to a server.
[1342] 3. The server analyzes the sign language data and identifies it as the sign for "thank you."
[1343] The server analyzes the received data using an image recognition algorithm and identifies it as the sign language for "thank you."
[1344] 4. The server retrieves background information and meaning of the sign language for "thank you" from the database and displays the retrieved information on the terminal.
[1345] The server retrieves information related to the sign language expression for "thank you" from its database. For example, this information may include details about how gratitude is expressed in Japanese culture. The retrieved information is then sent to the terminal as structured data (e.g., in JSON format).
[1346] 5. Display the information received by the device to the user.
[1347] The device visually displays the information it receives to the user. For example, it might display information such as, "The sign language for 'thank you' is used to express gratitude in Japanese."
[1348] Example of a prompt
[1349] The following is an example of a prompt statement for inputting a system description into the generating AI model.
[1350] Prompt message:
[1351] I would like to describe a system for learning sign language. This system helps users input sign language and understand its meaning and background information. Please explain the program's process in natural language, following the steps below.
[1352] procedure:
[1353] 1. The user inputs the sign language. The user either performs the sign language using their device's camera or uploads a pre-recorded video file of the sign language.
[1354] 2. The device sends sign language data to the server. The device saves the sign language video captured by the camera as digital data and sends it to the server via the internet.
[1355] 3. The server analyzes the sign language data. The server uses an image recognition algorithm to analyze what the sign language data means and compares it to an existing sign language dictionary.
[1356] 4. The server retrieves background information and meaning. The server queries the database using the identified sign language identifier to retrieve background information and meaning for that sign language.
[1357] 5. The server sends the information to the terminal. The server converts the acquired background information and meaning into structured data and sends it to the terminal.
[1358] 6. The terminal displays information. The terminal receives background information and meaning sent from the server and displays it to the user. Display formats include text, images, videos, etc.
[1359] Specific example:
[1360] Please explain the processing steps using the sign language for "thank you" as an example.
[1361] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1362] Specific flow of program processing
[1363] Processing steps
[1364] Step 1:
[1365] The user inputs sign language. The user either performs sign language using the camera on their device or uploads a pre-recorded video file of sign language. The input data is a video of sign language. This video of sign language becomes the raw data used for analysis in subsequent processing.
[1366] Step 2:
[1367] The device captures sign language data and saves it as digital data. The device's camera captures video frame by frame and saves this as digital data (e.g., MP4 format) to internal storage. The captured data is used for subsequent analysis.
[1368] Step 3:
[1369] The terminal transmits the stored sign language data to the server via a communication device. The terminal uses an internet connection to transmit the stored sign language data to the server. The input data is the captured sign language video data, and the output is the data transmitted to the server.
[1370] Step 4:
[1371] The server receives sign language data and analyzes it using a machine learning algorithm. The server then analyzes the received data using an image recognition algorithm (e.g., a Convolutional Neural Network). This analysis extracts the features of the sign language and converts them into an identifiable format. The input data is video data of sign language, and the output is the analyzed sign language data.
[1372] Step 5:
[1373] The server identifies a unique identifier based on the analyzed sign language data and retrieves background information and meaning from the data store. The server queries the database based on the analysis results to retrieve background information and meaning for the corresponding sign language. The input data is the analyzed sign language data, and the output is data related to background information and meaning.
[1374] Step 6:
[1375] The server converts the acquired information into structured data (e.g., JSON format) and sends it to the terminal. The server uses a data transmission function to send the acquired information to the terminal. The input data is background information and semantic data, and the output is the structured data sent to the terminal.
[1376] Step 7:
[1377] The terminal displays the information it receives to the user. The terminal's display device analyzes the received structured data and displays it visually to the user. Display formats include text, images, and videos. The input data is structured data, and the output is the information displayed to the user.
[1378] Detailed explanation of operation
[1379] Step 1:
[1380] Users perform the sign language for "thank you" using their camera. Users can also upload the recorded video file.
[1381] Step 2:
[1382] The device's camera captures video of the user performing sign language. This video is saved as digital data to the internal storage.
[1383] Step 3:
[1384] The device transmits the stored sign language video data to the server via the internet.
[1385] Step 4:
[1386] The server analyzes the received sign language video data using a Convolutional Neural Network. It extracts characteristic patterns from the sign language and converts them into specific identifiers.
[1387] Step 5:
[1388] The server uses the identified identifier to query the database and retrieve background information and meanings for the corresponding sign language.
[1389] Step 6:
[1390] The server converts the background information and meaning retrieved from the database into structured data in JSON format and sends it to the terminal.
[1391] Step 7:
[1392] The structured data received by the terminal is displayed on the display device, and information such as "The sign language for 'thank you' is used to express gratitude" is presented to the user.
[1393] (Application Example 1)
[1394] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1395] Conventional sign language learning systems not only teach sign language but also lack systems that enable its practical application in real-world work environments. In particular, it is difficult for deaf individuals to give work instructions to machines and robots using sign language in factories and manufacturing sites. In such environments, a work instruction system that utilizes sign language is needed.
[1396] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1397] In this invention, the server includes means for analyzing sign language data, means for obtaining background information and meaning related to the analyzed sign language data from a database, and means for identifying work instructions based on the obtained background information and meaning and transmitting them to a device. This enables reliable transmission of work instructions using sign language.
[1398] A "user" refers to a person who uses the system to input sign language and issue work instructions.
[1399] "Sign language" refers to a system of gestures and hand movements used by people with hearing impairments to communicate.
[1400] "Means" refers to the methods, devices, and components used to realize each function of a system.
[1401] "Data" refers to information including video recordings of sign language and the results of their analysis that the system processes.
[1402] "Server" refers to a central processing unit that performs sign language data analysis and generates work instructions.
[1403] "Analysis" refers to the processing of an algorithm that receives sign language data, understands its content, and identifies its meaning.
[1404] "Background information" refers to the cultural and historical contextual information associated with the analyzed sign language.
[1405] "Meaning" refers to the content or intention that sign language is meant to convey.
[1406] A "database" refers to a collection of information that stores sign language, its background information, and its meaning in an associated manner.
[1407] A "terminal" refers to a device used by a user to input sign language and receive and display the analysis results.
[1408] "Equipment" refers to a device that receives and executes work instructions generated from analyzed sign language data.
[1409] "Work instructions" refer to specific operational instructions that robots and other automated equipment should perform based on the analyzed sign language.
[1410] "Execution" refers to the physical actions or processes that a device performs based on the work instructions it receives.
[1411] Modes for carrying out the invention
[1412] The system that realizes this invention involves a user inputting sign language, analyzing that data to generate relevant work instructions, and transmitting those instructions to a device for execution. This system is mainly composed of the following hardware and software.
[1413] hardware
[1414] Camera module: This utilizes cameras attached to factory robots, as well as the built-in cameras of smartphones and tablets. This allows for the capture of user sign language.
[1415] Factory robots: Specifically, this includes automation equipment from companies such as KUKA, ABB, and Fanuc. These robots are responsible for receiving and executing work instructions.
[1416] software
[1417] Image recognition algorithm: This algorithm analyzes sign language data using tools such as TensorFlow and OpenCV. It extracts features from sign language and identifies what those signs mean.
[1418] Sign language database: MongoDB or PostgreSQL is used. This database stores the meaning and background information of sign language.
[1419] Server-side program: Uses Node.js or Python (Flask) to analyze sign language data and generate work instructions.
[1420] Frontend: JavaScript (React.js) is used to implement the user interface.
[1421] Specific examples of actions
[1422] 1. User inputs sign language: Users perform sign language towards a camera attached to the factory robot. Alternatively, they can input sign language from a smartphone or tablet. The sign language performed by the user mainly corresponds to work instructions such as "assemble" and "move".
[1423] Example prompt: "Perform the following sign language towards the camera and enter a work instruction: 'Move,' the robot will move."
[1424] 2. Data Transmission: The collected sign language data is sent to the server using WebSocket. The server divides the received sign language data into frames and prepares them for analysis.
[1425] 3. Data Analysis: An image recognition algorithm using TensorFlow is executed on the server side to analyze the received sign language data. This identifies the meaning of the sign language.
[1426] 4. Database query: The server uses the identified sign language identifier to retrieve the corresponding work instruction from a database such as MongoDB.
[1427] 5. Sending work instructions: The server sends the generated work instructions to the equipment. The equipment operates automatically based on the received work instructions.
[1428] Details of specific examples
[1429] For example, if a user performs the sign language for "move" while pointing it at the camera, the camera module captures the sign language. This data is sent to a server in real time and analyzed by TensorFlow. Based on the analysis results, the task instruction "move" is identified and retrieved from the database. The task instruction is then sent to the robot, which performs the specified movement.
[1430] This will enable people with hearing impairments to efficiently operate equipment in factories and manufacturing sites using sign language.
[1431] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1432] Step 1:
[1433] The user enters sign language.
[1434] The user performs sign language towards a camera attached to a factory robot or towards a smartphone or tablet. The entered sign language is captured as video and stored as digital data for use in the next processing step.
[1435] Input: User's sign language video
[1436] Output: Captured sign language video data
[1437] Step 2:
[1438] The device sends sign language data to the server.
[1439] The terminal (a robot or smartphone with a built-in camera module) divides the captured sign language video into frames in real time and sends them to the server using WebSocket.
[1440] Input: Captured sign language video data
[1441] Output: Sign language video frames sent to the server
[1442] Step 3:
[1443] The server analyzes the sign language data.
[1444] The server analyzes the received sign language data using a TensorFlow-based image recognition algorithm. It extracts each feature point of the sign language and performs analysis to determine what the sign language means.
[1445] Input: Sent sign language video frame
[1446] Output: Meaning of the analyzed sign language (identified sign language identifier)
[1447] Step 4:
[1448] The server retrieves background information and meaning from the database.
[1449] The server uses the identified sign language identifier to query a MongoDB or PostgreSQL database to retrieve background information, meaning, and corresponding work instructions for that sign language.
[1450] Input: Identified sign language identifier
[1451] Output: Background information, meaning, work instructions
[1452] Step 5:
[1453] The server sends work instructions to the equipment.
[1454] The server converts the acquired work instructions into structured data in JSON format and transmits it to equipment such as factory robots via the internet.
[1455] Input: Work instructions
[1456] Output: Work instructions sent to the device
[1457] Step 6:
[1458] The equipment performs tasks based on the work instructions it receives.
[1459] Factory automation equipment and robots analyze received work instructions and perform actions based on those instructions. For example, if the instruction is to move, the robot will move to the specified location.
[1460] Input: Received work instructions
[1461] Output: Performed actions (e.g., robot movement)
[1462] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1463] This invention relates to a system for learning sign language in a more accessible way, and in particular, a system that helps users input sign language and understand its background information, meaning, and even their emotions. The system consists of four main elements: the user, the terminal, the server, and the emotion engine.
[1464] Program Processing Overview
[1465] 1. The user enters the sign language.
[1466] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[1467] In this process, the device uses an emotion engine to acquire emotional data from the user's facial expressions and voice.
[1468] 2. The device sends sign language data and emotion data to the server.
[1469] The device stores sign language video captured by the camera and emotional data analyzed by the emotion engine as digital data.
[1470] Next, this sign language data and emotion data are sent to a server via the internet.
[1471] 3. The server analyzes the sign language data.
[1472] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Specifically, it extracts characteristic parts of the sign language and evaluates its similarity by comparing it with an existing sign language dictionary.
[1473] 4. The server identifies the sign language identifier.
[1474] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database.
[1475] 5. The server retrieves background information and meaning.
[1476] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[1477] 6. The server analyzes the emotional data.
[1478] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the user's emotional state is joyful when they input the sign language for "thank you."
[1479] 7. The server converts background information, semantic, and sentiment data into structured data.
[1480] The server converts the sign language background information, meaning, and emotional data into structured data and arranges it in a user-friendly format.
[1481] 8. The server sends the structured data to the terminal.
[1482] The server sends structured data regarding background information, meaning, and sentiment to the terminal.
[1483] 9. The device receives and displays information.
[1484] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, and videos, and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language.
[1485] Specific example
[1486] Example 1: Input of the sign language "thank you" and analysis of emotional data
[1487] The user performs the sign language for "thank you" while pointing it at the device's camera. The device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as the sign language for "thank you." It also uses an emotion engine to detect that the user's emotional state is joy. The server retrieves background information and meaning of the sign language for "thank you" from its database, adds the emotional state, and sends it to the device. The device then displays a message to the user stating something like, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1488] Example 2: Input of region-specific sign language and analysis of emotional data
[1489] The user inputs sign language specific to the Kansai region into the terminal. The terminal captures this sign language and sends it to the server along with emotion data. The server analyzes the sign language data and emotion data and recognizes that the sign language is specific to the Kansai region and that the user's emotional state is surprise or confusion. The server sends the regional background and cultural meaning of the sign language, as well as the emotional state, to the terminal. The terminal displays, "This sign language represents a greeting specific to the Kansai region, and the user's emotional state is 'surprise' or 'confusion'."
[1490] In this way, this system can improve the efficiency of sign language learning and the user's sense of familiarity by providing integrated background information, meaning, and the user's emotional state in sign language data.
[1491] The following describes the processing flow.
[1492] Step 1: The user enters sign language.
[1493] The user performs sign language gestures in front of the device's camera, or uploads a pre-recorded video file of sign language to the device.
[1494] The device captures sign language movements in real time using its camera.
[1495] Step 2: The device captures and saves sign language data and emotion data.
[1496] The device utilizes an emotion engine to analyze the user's emotional state in real time using facial recognition and voice analysis technologies.
[1497] These analysis results will be saved as digital data and managed together with the sign language data.
[1498] Step 3: The device sends sign language data and emotion data to the server.
[1499] The terminal transmits the stored sign language data and emotional data to the server via the internet. During this process, appropriate pre-processing, such as formatting and compression, is performed on the data.
[1500] Step 4: The server receives the sign language data.
[1501] The server receives sign language data and emotion data sent from the terminal.
[1502] Step 5: The server analyzes the sign language data.
[1503] The server analyzes the received sign language data and uses image recognition algorithms to identify what the sign language actions mean. For example, it extracts characteristic parts of the sign language actions and compares them with existing sign language dictionaries.
[1504] Step 6: The server identifies the sign language identifier.
[1505] The server identifies the corresponding sign language identifier based on the results of the image recognition algorithm. This identifier corresponds to a sign language entry in the database.
[1506] Step 7: The server retrieves background information and meaning.
[1507] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[1508] Step 8: The server analyzes the emotion data.
[1509] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it determines whether the user is exhibiting emotions such as "joy" or "surprise."
[1510] Step 9: The server structures all the data.
[1511] The server integrates background information, meaning, and emotional data from sign language and structures it into a format that is easy for users to understand.
[1512] Step 10: The server sends the structured data to the terminal.
[1513] The server sends integrated structured data to the terminal.
[1514] Step 11: The device receives and displays information.
[1515] The terminal receives structured data sent from the server and displays it to the user. The displayed content includes background information and meaning of sign language, as well as the user's emotional state. For example, it may be displayed visually in an easy-to-understand way using text, images, and videos.
[1516] Specific example
[1517] Example 1: Input of the sign language "thank you" and analysis of emotional data
[1518] 1. The user performs the sign language for "thank you" while pointing it at the device's camera.
[1519] 2. The device captures video and saves sign language data and emotion data analyzed by the emotion engine. The emotion engine detects the emotion of "joy" from the user's facial expressions and voice.
[1520] 3. The device sends sign language data and emotion data to the server.
[1521] 4. The server receives the sign language data.
[1522] 5. The server analyzes the sign language data using an image recognition algorithm.
[1523] 6. The server identifies the sign language identifier.
[1524] 7. The server retrieves background information and meaning of the sign language for "thank you" from the database.
[1525] 8. The server analyzes the emotion data and confirms that the user is experiencing the emotion of "joy."
[1526] 9. The server formats all data into structured data and sends it to the terminal.
[1527] 10. The device receives the data and displays the message, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1528] Example 2: Input of region-specific sign language and analysis of emotional data
[1529] 1. The user inputs sign language specific to the Kansai region into the device.
[1530] 2. The device captures sign language and sends it to the server along with emotion data. The emotion engine detects the emotion of "surprise" from the user's facial expressions.
[1531] 3. The server analyzes the sign language data and emotion data to confirm that it is a regionally specific sign language and to verify the emotional state.
[1532] 4. The server retrieves regional background and cultural significance from the database.
[1533] 5. The server integrates this data and sends it to the terminal.
[1534] 6. The device displays the message, "This sign language represents a greeting unique to the Kansai region, and the user's emotional state is 'surprise'."
[1535] Through this system, users can understand not only the background information and meaning of sign language, but also their own emotional state, making sign language learning more accessible.
[1536] (Example 2)
[1537] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1538] Conventional sign language learning systems have the drawback of not being able to simultaneously understand not only the background information and meaning of sign language, but also the emotional state of the user performing the sign language. Furthermore, there is a need to make learning sign language more accessible and effective by reflecting the user's emotional state when inputting sign language, but an effective system for this purpose has not existed.
[1539] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for analyzing transmitted sign language data and emotion data, means for obtaining background information and meaning related to the analyzed sign language data from a database, and means for evaluating the user's emotional state from the analyzed emotion data. This makes it possible to simultaneously grasp not only the background information and meaning of the sign language, but also the user's emotional state when performing the sign language.
[1540] A "user" refers to a person who uses the system to input sign language.
[1541] "Means of inputting sign language" refers to the interface (such as a camera or touchscreen) that a user uses to input sign language into a device.
[1542] "Sign language data" refers to information that digitally records the sign language actions entered by the user.
[1543] "Emotional data" refers to digital information that indicates the emotional state of a user, analyzed from their facial expressions, voice, and other data.
[1544] A "server" refers to a central computing device that performs data analysis, processing, and storage.
[1545] "Means of transmission" refers to the means of communication used to transfer data from a terminal to a server.
[1546] "Means of analysis" refers to algorithms and technologies (for example, image recognition algorithms and sentiment analysis technologies) that the server uses to process sign language data and sentiment data and analyze their content.
[1547] "Background information" refers to historical, cultural, and regional information related to sign language.
[1548] "Meaning" refers to the concept or intention expressed by the movements in sign language.
[1549] A "database" refers to an information management system that stores sign language data, background information, semantic data, and emotional data, making them accessible as needed.
[1550] A "specific identifier" refers to a code or label used to uniquely identify the type or meaning of sign language.
[1551] "Structured data" refers to data that follows a specific format and includes formalized data such as analyzed sign language information and emotional states.
[1552] A "terminal" refers to a device (such as a smartphone or tablet) used by a user to input sign language, receive data, and display it.
[1553] "Means of display" refers to the means by which a terminal visually presents analysis results and acquired information to the user.
[1554] This invention relates to a system for learning sign language in a more accessible way. The system helps users input sign language and understand its background information, meaning, and even their emotions. The system consists of four main elements: the user, the terminal, the server, and the emotion engine.
[1555] System Overview
[1556] 1. The user enters the sign language.
[1557] The user performs sign language actions in front of the device's camera. For example, to input the sign language for "thank you," the user lightly presses their palms in front of their chest. Alternatively, they can upload a pre-recorded video file of sign language to the device. In this case, the device's camera captures the user's sign language movements, and an emotion engine is used to obtain emotional data from the user's facial expressions and voice.
[1558] 2. The device sends sign language data and emotion data to the server.
[1559] The device captures sign language video with its camera and saves it as digital data. It also collects emotional data analyzed by an emotion engine. The captured sign language data and emotional data are sent together as a data packet to the server via the internet. Examples of emotion engines that can be used include "OpenCV" and "Microsoft Azure Face API".
[1560] 3. The server analyzes the sign language data.
[1561] The server receives sign language data sent from the terminal and uses an image recognition algorithm to analyze what the sign language data means. Examples of image recognition algorithms used include "TensorFlow" and "PyTorch". It extracts characteristic parts of the sign language and evaluates similarity by comparing them with existing sign language dictionaries.
[1562] 4. The server identifies the sign language identifier.
[1563] Based on the results of the image recognition algorithm, the server identifies the identifier for the corresponding sign language. This identifier corresponds to a sign language entry in the database. For example, the sign language for "thank you" is identified as identifier "S001".
[1564] 5. The server retrieves background information and meaning.
[1565] The server uses the identified sign language identifier to retrieve background information about the sign language (such as its origin, history, and cultural background) and its meaning from the database.
[1566] 6. The server analyzes the emotional data.
[1567] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the emotional state of a user when they input the sign language for "thank you" is "joy." Possible emotional analysis platforms used include "IBM Watson" and "Amazon Rekognition."
[1568] 7. The server converts background information, semantic, and sentiment data into structured data.
[1569] The server converts the sign language background information, meaning, and emotional data into structured data and arranges it in a user-friendly format. For example, it converts it into a structured data format such as JSON.
[1570] 8. The server sends the structured data to the terminal.
[1571] The server sends structured data containing background information, semantic, and sentiment data to the terminal. Communication protocols used include HTTP and WebSocket.
[1572] 9. The device receives and displays information.
[1573] The terminal receives structured data sent from the server and displays it to the user. The display format includes text, images, and videos, and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language. For example, the terminal screen may display the message, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1574] Specific example
[1575] Example 1: Input of the sign language "thank you" and analysis of emotional data
[1576] The user performs the sign language for "thank you" towards the device's camera, and the device captures the video and sends it to the server as sign language data. The server analyzes the sign language data and identifies it as identifier "S001". It also uses an emotion engine to detect that the user's emotional state is "joy". The server retrieves background information and meaning of the "thank you" sign language from the database, adds the emotional state, and sends it to the device. The device displays to the user, "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1577] Example 2: Input of region-specific sign language and analysis of emotional data
[1578] The user inputs a sign language specific to the Kansai region into the terminal, and the terminal captures this sign language and sends it to the server along with emotion data. The server analyzes the sign language data and emotion data and recognizes that the sign language is specific to the Kansai region and that the user's emotional state is "surprise" or "confusion". The server sends the regional background and cultural meaning of the sign language, as well as the emotional state, to the terminal. The terminal displays, "This sign language expresses a greeting specific to the Kansai region, and the user's emotional state is 'surprise' or 'confusion'."
[1579] Example of a prompt
[1580] 1. "Please analyze the emotions expressed when performing sign language and provide this information along with the background context."
[1581] 2. What system can analyze the sign language for "thank you" and display the emotional state?
[1582] In this way, this system can improve the efficiency of sign language learning and the user's sense of familiarity by providing integrated background information, meaning, and the user's emotional state in sign language data.
[1583] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1584] Step 1:
[1585] The user enters sign language.
[1586] The user performs sign language actions in front of the device's camera. For example, to sign "thank you," the user lightly presses their palms in front of their chest. Alternatively, they can upload a pre-recorded sign language video file to the device. Input can be real-time sign language video or uploaded video files. Output is the device acquiring this sign language video data.
[1587] Step 2:
[1588] The device sends sign language data and emotion data to the server.
[1589] The device stores sign language video captured by its camera as digital data and uses an emotion engine to analyze emotional data from the user's facial expressions and voice. For example, the emotion engine uses facial recognition technology to analyze the degree of the user's smile and the tone of their voice. The input consists of the captured sign language video data and the results of the emotion analysis. As output, this data is sent to the server as data packets.
[1590] Step 3:
[1591] The server analyzes the sign language data.
[1592] The server inputs the received sign language data into an image recognition algorithm, extracts the characteristic parts of the sign language, and begins analysis. Examples of algorithms used include "TensorFlow" and "PyTorch." The input is sign language video data. The output is a sign language feature vector.
[1593] Step 4:
[1594] The server identifies the sign language identifier.
[1595] The server compares the generated feature vectors with existing sign language dictionaries to find the most similar sign language data. Each identified sign language data is assigned a unique identifier. The input consists of feature vectors and a sign language dictionary. The output is a sign language identifier (e.g., "S001").
[1596] Step 5:
[1597] The server retrieves background information and meaning.
[1598] The server uses the identified sign language identifier to retrieve background information (such as the region of origin, history, and cultural background) and its meaning from the database. The input is the sign language identifier. The output is the background information and meaning of the sign language.
[1599] Step 6:
[1600] The server analyzes the emotional data.
[1601] The server analyzes the received emotional data and evaluates the user's emotional state. For example, it analyzes whether the user's emotional state when they sign "thank you" is "joy." The input is emotional data, and the output is an evaluation of the user's emotional state.
[1602] Step 7:
[1603] The server converts background information, semantic, and sentiment data into structured data.
[1604] The server converts the sign language background information, meaning, and analyzed sentiment data into structured data (e.g., JSON format) and arranges it in a user-friendly format. Input includes background information, meaning, and sentiment data. The output is structured data.
[1605] Step 8:
[1606] The server sends structured data to the terminal.
[1607] The server sends structured data to the terminal. Communication protocols used include "HTTP" and "WebSocket". The input is structured data. The output is this data sent to the terminal.
[1608] Step 9:
[1609] The device receives and displays information.
[1610] The terminal receives structured data sent from the server and passes it to the UI for display to the user. The display format includes text, images, videos, etc., and is provided in a way that makes it easy for the user to understand the background information, meaning, and emotional state of the sign language. The input is structured data. The output is information that is displayed to the user. For example, the terminal screen might display "The sign language for 'thank you' is used to express gratitude in Japanese, and the user's emotional state is 'joy'."
[1611] (Application Example 2)
[1612] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1613] This invention relates to a system for improving the ease of use and usability of electronic payment systems for users who use sign language. Conventional electronic payment systems are based on voice or text-based input, making them difficult to operate for users who use sign language. Furthermore, they lacked the function to analyze the user's emotional state and provide appropriate feedback accordingly. Therefore, there is a need to provide a system that integrates sign language input and emotion analysis to enable users who use sign language to comfortably make electronic payments.
[1614] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing input sign language data and user emotion data, means for obtaining background information and meaning related to the analyzed sign language data, as well as user emotion information, from a database, and means for transmitting the obtained background information, meaning, and emotion information to the terminal. As a result, when a user who uses sign language uses the electronic payment system, not only is the background information and meaning of the sign language displayed in an easy-to-understand manner, but appropriate feedback according to the user's emotional state is also provided.
[1615] A "user" refers to an individual who operates the system and inputs sign language and emotional data.
[1616] "Sign language data" refers to information that digitizes actions entered by users using sign language.
[1617] "Emotional data" refers to digitized information that represents the emotional state of a user, analyzed from their facial expressions and voice.
[1618] A "terminal" refers to a device used by users to input sign language and emotional data and send that information to a server.
[1619] A "server" refers to a device that analyzes sign language data and emotional data, retrieves the results from a database, and transmits them to a terminal.
[1620] An "emotion recognition algorithm" refers to a computational method for analyzing a user's emotional state from their facial expressions and voice.
[1621] An "image recognition algorithm" refers to a computational method for analyzing sign language data and identifying sign language movements.
[1622] An "identifier" refers to a specific feature or attribute associated with the analyzed sign language data and sentiment data.
[1623] "Background information" refers to information related to sign language, such as its place of origin, history, and cultural background.
[1624] "Meaning" refers to the concept or intention represented by a particular sign language action.
[1625] An "electronic payment system" refers to a system that allows users to conduct online transactions using digital currencies, credits, etc.
[1626] This invention provides a system that enables users who use sign language to smoothly utilize electronic payment systems. This system is implemented with the following configuration and means.
[1627] First, users input sign language using a device such as a smartphone. The device has a built-in camera, which can capture the user's sign language movements as video. Users can also input voice, which allows for the acquisition of facial expressions and vocal information. This data is collected as "sign language data" and "emotional data."
[1628] Next, the terminal sends the collected sign language data and emotion data to the server. An internet connection is used for transmission. The server implements image recognition algorithms and emotion recognition algorithms for analyzing the sign language and emotion data. Specifically, software such as TensorFlow and OpenCV are used. This analyzes the sign language movements, facial expressions, and voice to identify the user's intended meaning.
[1629] When the server analyzes sign language data and emotion data, it retrieves background information and meaning from a database based on the content. This database stores information such as the region of origin, history, and cultural background of the sign language, as well as the meaning of the sign language. The emotion data also includes the user's emotional state, which is also analyzed by the server. Based on the results of this analysis, the server retrieves background information and meaning of the sign language, as well as the user's emotional state, from the database along with a specific identifier.
[1630] Next, the server sends the acquired background information, meaning, and emotional information to the terminal. The terminal displays the received information to the user in an easy-to-understand format. The display format includes text, images, videos, etc., and is provided in a way that is easy for the user to understand. This ensures that when a user of sign language uses the electronic payment system, not only is the background information and meaning of the sign language clearly displayed, but appropriate feedback corresponding to the emotions the user is feeling is also provided at the same time.
[1631] As a concrete example, consider a case where a user performs the sign language for "thank you" in front of their smartphone camera. The device captures the video and sends it to the server. The server analyzes the sign language data to identify it as the sign language for "thank you," and uses an emotion recognition algorithm to detect that the user's emotional state is joy. The server retrieves background information and meaning of the "thank you" sign language, as well as the user's emotional state, from a database and sends it to the device. The device then displays a message to the user stating something like, "'Thank you' is a sign language expression used to express gratitude, and the user's emotional state is 'joy'."
[1632] Examples of prompt messages are as follows:
[1633] example:
[1634] Please generate a Python program that analyzes the following video and recognizes the meaning of the sign language and the user's emotions. The sign language should be returned as Unicode, and the emotions as text.
[1635] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1636] Step 1:
[1637] Users input sign language using their smartphones. The device's camera is used as the input device, capturing the user's sign language movements in real time. The user's facial expressions and voice are also recorded simultaneously. This data is collected as "sign language data" and "emotional data."
[1638] Input: User's sign language gestures, facial expressions, and voice.
[1639] Output: Sign language data, emotion data (captured video and audio)
[1640] Step 2:
[1641] The terminal collects sign language data and emotional data and sends it to the server. An internet connection is used for transmission. Furthermore, the data undergoes necessary preprocessing (e.g., frame resizing and normalization) before being sent to the server.
[1642] Input: Sign language data, emotion data
[1643] Output: Digital data sent to the server
[1644] Step 3:
[1645] The server analyzes sign language data and emotion data. Sign language data is analyzed using an image recognition algorithm (e.g., a TensorFlow model) to identify the gestures and meanings of the signs. Emotion data is analyzed using an emotion recognition algorithm (e.g., a TensorFlow / Keras model) to identify the user's emotional state.
[1646] Input: Sign language data, emotion data
[1647] Output: Analysis results (sign language identification information, emotional state)
[1648] Step 4:
[1649] Based on the analysis results, the server retrieves background information and meanings of sign language, as well as user emotional information, from the database. The database stores the region of origin, history, cultural background, and meaning of each sign language. Appropriate feedback information tailored to the user's emotional state is also retrieved.
[1650] Input: Analysis results (sign language identification information, emotional state)
[1651] Output: Sign language background information, semantic information, and emotional information
[1652] Step 5:
[1653] The server transmits background information, semantic information, and emotional information acquired by the server to the terminal. The data is structured and presented in a user-friendly format. An internet connection is used for transmission.
[1654] Input: Background information, meaning, sentiment information
[1655] Output: Structured data sent to the terminal
[1656] Step 6:
[1657] The device displays received information to the user. The display format includes text, images, and videos, and is provided in a way that allows the user to easily understand the background information, meaning, and emotional state of the sign language.
[1658] Input: Received structured data
[1659] Output: Information displayed to the user (background information, meaning, emotional state)
[1660] Through these steps, users can smoothly utilize the electronic payment system using sign language, and in the process, receive appropriate feedback that reflects the background information and meaning of the sign language, as well as their own emotional state.
[1661] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1662] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1663] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1664] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1665] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1666] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1667] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1668] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1669] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1670] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1671] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1672] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1673] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1674] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1675] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1676] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1677] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1678] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1679] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1680] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1681] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1682] The following is further disclosed regarding the embodiments described above.
[1683] (Claim 1)
[1684] A means for users to input sign language,
[1685] A means for sending the input sign language data to the server,
[1686] A means for analyzing the transmitted sign language data,
[1687] A means of obtaining background information and meaning related to the analyzed sign language data from a database,
[1688] A means of transmitting the acquired background information and meaning to the terminal,
[1689] A means of displaying information received by the terminal to the user,
[1690] A system that includes this.
[1691] (Claim 2)
[1692] The system according to claim 1, which uses an image recognition algorithm to analyze sign language data.
[1693] (Claim 3)
[1694] The system according to claim 1, which uses a specific identifier to associate background information and meaning of sign language.
[1695]
[1696] "Example 1"
[1697] (Claim 1)
[1698] A means including an input device in which the user inputs sign language,
[1699] A means for transmitting input sign language data to a server via a communication device,
[1700] A means of analyzing transmitted sign language data using a machine learning algorithm,
[1701] A means of obtaining background information and meaning related to the analyzed sign language data from a data store,
[1702] A means of sending the acquired background information and meaning to the terminal,
[1703] A means for displaying information received by a terminal on a display device,
[1704] A system that includes this.
[1705] (Claim 2)
[1706] The system according to claim 1, which uses an image recognition algorithm to analyze sign language data.
[1707] (Claim 3)
[1708] The system according to claim 1, which uses a unique identifier to associate background information and meaning of sign language.
[1709] "Application Example 1"
[1710] (Claim 1)
[1711] A means for users to input sign language,
[1712] A means for sending the input sign language data to the server,
[1713] A means for analyzing the transmitted sign language data,
[1714] A means of obtaining background information and meaning related to the analyzed sign language data from a database,
[1715] A means of transmitting the acquired background information and meaning to the terminal,
[1716] A means of displaying information received by the terminal to the user,
[1717] A means for identifying work instructions and transmitting those work instructions to a device,
[1718] A means by which the equipment performs work based on the work instructions it receives,
[1719] A system that includes this.
[1720] (Claim 2)
[1721] The system according to claim 1, which uses an image recognition algorithm to analyze sign language data.
[1722] (Claim 3)
[1723] The system according to claim 1, which uses a specific identifier to associate background information and meaning of sign language.
[1724] "Example 2 of combining an emotion engine"
[1725] (Claim 1)
[1726] A means for users to input sign language,
[1727] A means for transmitting input sign language data and emotional data to a server,
[1728] A means for analyzing transmitted sign language data and emotional data,
[1729] A means of obtaining background information and meaning related to the analyzed sign language data from a database,
[1730] A means of evaluating the user's emotional state from analyzed emotional data,
[1731] A means for converting acquired background information, meaning, and emotional states into structured data,
[1732] A means of transmitting structured data to a terminal,
[1733] A means of displaying structured data received by the terminal to the user,
[1734] A system that includes this.
[1735] (Claim 2)
[1736] The system according to claim 1, which uses an image recognition algorithm to analyze sign language data.
[1737] (Claim 3)
[1738] The system according to claim 1, which uses a specific identifier to associate background information and meaning of sign language.
[1739] "Application example 2 when combining with an emotional engine"
[1740] (Claim 1)
[1741] A means for users to input sign language,
[1742] A means for transmitting the input sign language data and user emotion data to a server,
[1743] A means for analyzing transmitted sign language data and emotional data,
[1744] A means of obtaining background information, meaning, and user emotion information related to the analyzed sign language data from a database,
[1745] A means for transmitting acquired background information, semantic information, and emotional information to a terminal,
[1746] A means of displaying information received by the terminal to the user,
[1747] A system that includes this.
[1748] (Claim 2)
[1749] The system according to claim 1, which uses an image recognition algorithm and an emotion recognition algorithm to analyze sign language data and user emotions.
[1750] (Claim 3)
[1751] The system according to claim 1, which uses specific identifiers to associate background information, meaning, and emotional information of sign language. [Explanation of Symbols]
[1752] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for users to input sign language, A means for sending the input sign language data to the server, A means for analyzing the transmitted sign language data, A means of obtaining background information and meaning related to the analyzed sign language data from a database, A means of transmitting the acquired background information and meaning to the terminal, A means of displaying information received by the terminal to the user, A system that includes this.
2. The system according to claim 1, which uses an image recognition algorithm to analyze sign language data.
3. The system according to claim 1, which uses a specific identifier to associate background information and meaning of sign language.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A