system
A system for capturing and analyzing sign language movements to generate sentence candidates addresses the inefficiencies of current technologies, enabling efficient and intuitive communication by reducing user burden.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Current technologies require significant effort and time for sign language recognition and text generation, imposing a heavy burden on sign language users, and there is a need for an efficient and intuitive communication method.
A system that captures sign language movements, analyzes the data to recognize specific words, generates sentence candidates, and allows users to select from these candidates, reducing the burden through capture, analysis, generation, and presentation means.
The system enables efficient and intuitive communication by streamlining the process of converting sign language into text, allowing users to easily generate and convey messages.
Smart Images

Figure 2026062143000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] For users who communicate using sign language, there is a problem that there is no method for efficiently generating text and conveying the text to others. Current technologies require a lot of effort and time for sign language recognition and text generation, imposing a heavy burden on sign language users. There is a need to solve this problem and provide an efficient and intuitive communication method using sign language.
Means for Solving the Problems
[0005] This invention provides a system that captures a user's sign language movements, analyzes the captured data, and recognizes specific words. Furthermore, it includes means for generating sentence candidates based on the recognized words and presenting these candidates to the user, enabling the user to select from the generated sentence candidates. By including capture means, analysis means, generation means, presentation means, and selection means, this invention reduces the burden on sign language users and enables efficient and intuitive communication.
[0006] A "user" refers to a person who operates the system and inputs sign language actions.
[0007] "Motion data" refers to information captured by cameras and sensors of the user's sign language movements.
[0008] "Means of capturing" refers to devices such as cameras and sensors used to record user actions.
[0009] "Means of analysis" refers to software or algorithms used to process captured motion data and recognize specific words.
[0010] "Specific words" refer to each element of sign language recognized from the analyzed motion data (e.g., "subject," "predicate," "object").
[0011] "Generating means" refers to algorithms or programs used to create sentence candidates based on recognized words.
[0012] "Sentence suggestions" refer to multiple variations of sentences that the user can select from, generated from recognized words.
[0013] "Means of presentation" refers to displays or screens used to visually show the generated text candidates to the user.
[0014] "Means of selection" refers to the interface or operation methods that allow the user to choose their desired sentence from the presented sentence options.
[0015] The "system" refers to an aggregate of hardware and software necessary to execute a series of processes of capture, analysis, generation, presentation, and selection.
Brief Description of the Drawings
[0016] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing apparatus and a robot according to the fourth embodiment. [Figure 9] [[ID=3*]]It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0017] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0018] First, the terms used in the following description will be explained.
[0019] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0020] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0021] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0022] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0024] [First Embodiment]
[0025] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0026] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0027] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0028] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0029] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0031] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0032] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0033] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0034] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0035] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0036] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0037] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate and communicate text using sign language.
[0038] Users perform sign language actions through a device equipped with a camera and sensors. The device uses the camera to capture these actions in real time. The camera continuously acquires video data and transmits it to a server.
[0039] The server analyzes the received video data and converts sign language movements into specific words. This uses image processing techniques and machine learning algorithms. Specifically, the server recognizes the shape and movement patterns of the hands and identifies the corresponding words based on a pre-trained model.
[0040] Next, the server generates sentence candidates based on the recognized words. Natural language processing (NLP) techniques are used here. The server generates multiple sentence candidates based on the recognized word combinations and lists them.
[0041] The generated sentence candidates are sent to the terminal, which then visually presents them to the user. Specifically, the candidate sentences are displayed on the screen, allowing the user to select one. The user uses finger movements or touch gestures to select the desired sentence from the presented candidates.
[0042] The device detects the user's selection and displays the final selected sentence. This allows the user to easily confirm the message they want to convey.
[0043] For example, if a user performs the sign language actions for "I," "go," and "school," the server recognizes these as specific words and generates several sentence options, such as "I go to school," "It is I who goes to school," and "The place I go to is school." The user then selects their desired sentence from these options, and the terminal displays the selected sentence as the final output.
[0044] In this way, the present invention is a system that can streamline communication using sign language and significantly reduce the burden on sign language users.
[0045] The following describes the processing flow.
[0046] Step 1:
[0047] The user performs sign language actions. A device equipped with a camera and sensors captures these sign language actions in real time. Specifically, the device continuously acquires video data, accurately capturing the shape and movement of the hand.
[0048] Step 2:
[0049] The terminal sends the acquired video data to the server. The server receives this video data and begins analysis. At this stage, data preprocessing such as noise reduction and color correction may be performed.
[0050] Step 3:
[0051] The server analyzes video data and converts sign language movements into specific words. The server uses image processing technology to analyze the contours and movements of the hands and maps the recognized movements to corresponding words through machine learning algorithms. This identifies each element of the sign language (e.g., subject, predicate, object, etc.).
[0052] Step 4:
[0053] The server uses natural language processing (NLP) techniques to generate sentence candidates based on the words it recognizes. For example, from the words "I," "go," and "school," multiple sentences such as "I go to school" and "It is I who go to school" are created.
[0054] Step 5:
[0055] The server sends generated sentence candidates to the terminal, which then visually presents them to the user. Specifically, multiple candidate sentences are displayed on the screen in a list format. Each sentence candidate is identified using a number or highlighting.
[0056] Step 6:
[0057] The user selects their desired text from a list of suggested texts. The user makes the selection using finger movements or touch gestures. The device captures the user's selection actions and sends that information to the server.
[0058] Step 7:
[0059] The server receives the user's selection information and finally confirms the selected text. The selected text is then sent from the server to the terminal.
[0060] Step 8:
[0061] The device visually displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others.
[0062] Through the processing steps described above, the present invention realizes efficient and intuitive text generation and communication using sign language.
[0063] (Example 1)
[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0065] While modern communication largely relies on text-based technologies, their use can be limited for users who primarily rely on sign language, such as those with hearing impairments. Therefore, there is a need for systems that can efficiently and intuitively generate text using sign language and transmit it directly to others. However, existing technologies have been insufficient in accurately recognizing and transcribing sign language movements, hindering smooth communication for sign language users. In particular, challenges existed in the accurate capture and analysis of sign language movements, and the subsequent natural text generation process.
[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0067] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, and means for generating sentence candidates based on the recognized words. This enables the real-time capture of the user's sign language actions, transmission of the action data to the server, and analysis of the sign language actions using image processing technology and machine learning algorithms. Furthermore, by utilizing natural language processing technology to generate multiple sentence candidates and presenting them visually to the user, the user can easily select their desired sentence and confirm the final sentence. This enables sign language users, including those with hearing impairments, to communicate more efficiently and intuitively.
[0068] "Means for capturing user actions" refer to devices or methods for recording a user's sign language movements in real time, and generally involve the use of cameras or sensors.
[0069] "Means for analyzing captured motion data and recognizing specific words" refers to technologies for analyzing acquired video data of sign language movements and identifying specific words based on that analysis, and includes image processing technologies and machine learning algorithms.
[0070] "A means of generating sentence candidates based on recognized words" refers to a system that automatically generates appropriate sentences by combining words recognized through analysis, and utilizes natural language processing technology.
[0071] "Means for presenting generated text candidates to the user" refers to devices or methods for displaying text candidates in a way that the user can easily review, and primarily uses displays or monitors.
[0072] "Means for the user to select generated sentence candidates" refers to an interface for the user to choose their preferred sentence from among several presented candidates, and includes touchscreens, pointing devices, and the like.
[0073] "Means for displaying selected text" refers to devices or methods for visually showing the text ultimately selected by the user, and displays or screens are used for this purpose.
[0074] "Methods of using cameras and sensors for capture" refers to technologies that utilize cameras and various sensors to record detailed sign language movements.
[0075] "Image processing techniques for identifying hand shape and movement patterns" are methods for characterizing hand shape and movement by analyzing data acquired from cameras and sensors, and utilize image processing libraries, etc.
[0076] "Methods for recognizing words using trained machine learning algorithms" refers to technologies that automatically identify corresponding words from sign language gestures using machine learning models that have been trained on a large amount of data in advance.
[0077] "Methods for generating sentence candidates using natural language processing technology" refers to methods that utilize machine learning models to generate multiple natural-sounding sentences from recognized word combinations.
[0078] "Means for a terminal to send video data acquired by its camera to a server" refers to a technology that allows a terminal to send video data captured in real time to a server via communication means such as the internet.
[0079] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate and communicate text using sign language. Specific embodiments are described below.
[0080] Hardware and software configuration
[0081] Hardware:
[0082] 1. Camera: Used to capture the user's sign language movements. Generally, a high-resolution camera is used.
[0083] 2. Sensors: Used to help accurately determine the position of hands and fingers. Depth sensors are examples of such sensors.
[0084] 3. Terminal: A device used to process captured video data and display the results to the user. Smartphones and tablets are typical examples.
[0085] 4. Display: Built into the device and used to visually present text suggestions to the user.
[0086] software:
[0087] 1. Image processing technology: Using libraries such as OpenCV, we extract the shape and movement characteristics of the hand from video data captured by the camera.
[0088] 2. Machine Learning Algorithms: Using tools such as TENSORFLOW® or PyTorch, we construct and apply a model to convert sign language actions into specific words based on the extracted features.
[0089] 3. Natural Language Processing (NLP): Using generative AI models such as GPT and BERT, natural-sounding sentence candidates are generated based on recognized words.
[0090] Specific description of the system's operation
[0091] The user stands in front of the device's camera to perform sign language actions. The device's camera captures the user's sign language actions in real time. The video data acquired by the camera is transmitted to a server via the internet. The server analyzes the received video data and uses image processing technology to identify hand shapes and movement patterns. Next, a machine learning algorithm is applied to convert the identified sign language actions into corresponding words.
[0092] Next, natural language processing technology is used to generate multiple sentence candidates based on the recognized words. The generated sentence candidates are sent from the server to the terminal and displayed on the terminal's screen. The user selects the desired sentence from the presented candidates using touch gestures or other methods. Finally, the selected sentence is displayed on the terminal's screen, allowing the user to show the content they want to communicate to others.
[0093] Specific example
[0094] As an example, consider a case where a user performs the sign language actions for "I," "go," and "school." The device's camera captures these actions and sends the data to a server. The server analyzes the video data and identifies the words "I," "go," and "school." Next, using natural language processing technology, it generates several sentence options, such as "I go to school," "It is I who goes to school," and "The place I go to is school." The generated sentence options are sent to the device and displayed on the screen. The user selects their desired sentence, and the final sentence is displayed on the screen.
[0095] Example of a prompt
[0096] Example: "Please write a prompt for a system that analyzes sign language movements and generates multiple sentence options based on them."
[0097] As described above, the system of the present invention can significantly improve communication using sign language and reduce the communication burden on sign language users.
[0098] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0099] Step 1: Capture the user's sign language movements
[0100] Input: User's sign language gestures
[0101] Specific operation: The user stands in front of the device's camera to perform sign language actions. The device's camera captures the sign language actions in real time and continuously acquires video data.
[0102] Output: Acquired video data
[0103] Step 2: Sending video data
[0104] Input: Acquired video data
[0105] Specific operation: The device sends the captured video data to the server. The data is streamed to the server in real time over the internet.
[0106] Output: Video data sent to the server
[0107] Step 3: Analyzing video data
[0108] Input: Video data sent to the server
[0109] Specific operation: The server analyzes the received video data. First, it uses image processing techniques (such as OpenCV) to extract features of the hand's shape and movement. Afterwards, it identifies feature quantities (hand shape, position, movement pattern, etc.).
[0110] Output: Feature data
[0111] Step 4: Word Recognition
[0112] Input: Feature data
[0113] Specific operation: The server applies machine learning algorithms (TensorFlow or PyTorch) based on the extracted features to convert sign language actions into corresponding words. A pre-trained model identifies the appropriate word from the sign language action.
[0114] Output: Recognized words
[0115] Step 5: Generating text candidates
[0116] Input: Recognized word
[0117] Specific operation: The server generates multiple sentence candidates using natural language processing (NLP) techniques (GPT and BERT) based on recognized words. It then automatically generates appropriate sentence combinations using a generative AI model.
[0118] Output: Sentence suggestion list
[0119] Step 6: Send the suggested message
[0120] Input: Sentence suggestion list
[0121] Specific operation: The server sends a list of generated sentence suggestions to the terminal. The data is sent to the terminal via the network.
[0122] Output: List of suggested sentences sent to the terminal
[0123] Step 7: Display candidate sentences
[0124] Input: List of suggested sentences sent to the terminal
[0125] Specific operation: The device displays received text suggestions on its screen. It displays a list so that the user can visually see multiple text suggestions.
[0126] Output: Suggested sentences displayed on the screen
[0127] Step 8: User Selection
[0128] Input: Suggested sentences displayed on the screen
[0129] Specific operation: The user selects the desired sentence from the presented sentence options using finger movements or touch operations.
[0130] Output: Selected text
[0131] Step 9: Display the final sentence
[0132] Input: Selected text
[0133] Specific operation: The device displays the final sentence selected by the user on its screen. This allows the user to show others the message they ultimately wanted to convey.
[0134] Output: The final sentence displayed on the screen
[0135] (Application Example 1)
[0136] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0137] Traditional factory robot operation typically involves keyboards or touch panels, limiting intuitive control methods. This posed a problem, particularly for workers whose primary means of communication is sign language, as the operation was complex and difficult to understand. Furthermore, the lack of sufficient means to improve work efficiency was also a challenge.
[0138] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0139] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words, means for presenting the generated sentence candidates to the user, means for the user to select a generated sentence candidate, means for the user to give instructions for robot operation using sign language, and means for robot operation to execute the recognized sign language instructions. This makes it possible to operate the robot intuitively and efficiently using sign language.
[0140] "Means for capturing user movements" refers to devices that use cameras or sensors to capture the user's body movements and acquire those movements as digital data.
[0141] "A means of analyzing captured motion data and recognizing specific words" refers to the process of analyzing acquired digital data and identifying words that correspond to specific sign language or actions intended by the user.
[0142] "Methods for generating sentence candidates based on recognized words" refers to algorithms that combine words identified through analysis to generate candidates that form natural-sounding sentences.
[0143] "Means of presenting generated sentence candidates to the user" refers to an interface that visually displays multiple generated sentences as candidates on a display or screen, allowing the user to select one.
[0144] "Means for the user to select generated sentence candidates" refers to the means by which the user selects the desired sentence from among multiple displayed sentence candidates.
[0145] "A means by which a user can give instructions to operate a robot using sign language" refers to an interface for transmitting operating instructions to a robot or machine through sign language actions.
[0146] A "robot operating means for executing instructions based on recognized sign language" is a control device that allows a robot to perform specific operations based on instructions recognized as sign language actions.
[0147] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate text using sign language and to operate robots and give instructions in a factory setting.
[0148] hardware
[0149] The server and terminal include the following hardware:
[0150] Camera (e.g., Logitech C920): Captures the user's sign language movements in real time.
[0151] Robot control device: A device for controlling a robot based on recognized sign language movements.
[0152] software
[0153] The following software will be used:
[0154] Python: A programming language used for the overall system implementation.
[0155] OpenCV (CV2): A library for capturing and pre-processing camera footage.
[0156] Keras: A machine learning framework for running sign language gesture recognition models.
[0157] Natural language processing libraries (e.g., spaCy, GPT model): These are libraries for generating text candidates based on sign language actions.
[0158] System operation
[0159] The server receives video data of sign language movements transmitted from the user's camera. The video data is first preprocessed using OpenCV and then fed into a Keras model. The Keras model uses a pre-trained machine learning algorithm to convert the sign language movements into corresponding words.
[0160] The converted words are then transformed into multiple sentence candidates using a natural language processing library (e.g., spaCy or the GPT model). These sentence candidates are sent from the server to the terminal and displayed visually on the terminal's screen.
[0161] The user selects their desired sentence from several sentence options displayed on the screen. This selection is made through touch operations or finger movements. The final selected sentence is displayed on the device and then sent to the robot.
[0162] The robot performs specific actions based on the selected sentence. For example, if the user performs the sign language action for "pick up the part and move it," the robot will receive instructions such as "pick up the part and move it 5 meters."
[0163] Specific example
[0164] When a worker performs a sign language gesture for "pick up the part and move it," the server analyzes it and recognizes it as a corresponding word. The server then uses natural language processing technology to generate sentence options such as "pick up the part and move it 5 meters" or "take the part and move it to the designated location." These options are sent to the terminal, and the user selects the one sentence they prefer. Finally, the robot performs the action based on the selected instruction.
[0165] Example of a prompt
[0166] Example of sign language recognition results for use in analyzing sign language movements and generating text: Please generate phrases corresponding to the sign language words "part," "pick up," and "move." Possible instruction sentences would be something like, "Pick up the part and move 5 meters."
[0167] This invention enables intuitive and efficient operation and instruction using sign language in factories. Furthermore, it makes operation easier and improves work efficiency for workers who primarily use sign language for communication.
[0168] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0169] Step 1:
[0170] The server receives video data of sign language movements transmitted from the user's camera in real time. This video data is data captured by the camera of the sign language movements performed by the user. The data input is a video stream from the camera, and the output is the frame data of that video. Specifically, the camera continuously captures the user's hand movements and sends each frame to the server as digital data.
[0171] Step 2:
[0172] The server analyzes the received video data. OpenCV is used for preprocessing the video. In this step, the necessary sign language movements are extracted from the video, and noise is removed. The input for this process is frame data, and the output is preprocessed image data of the sign language movements. Specifically, this involves image smoothing and edge detection to enhance the shape of the hands.
[0173] Step 3:
[0174] The server supplies pre-processed video data to a Keras model, which runs a machine learning algorithm that converts sign language gestures into corresponding words. In this step, the input is pre-processed image data, and the output is the recognized words. Specifically, the model analyzes certain features in the image and maps them to pre-trained words.
[0175] Step 4:
[0176] The server generates multiple sentence candidates using a natural language processing library based on the recognized words. In this step, the input is the recognized words, and the output is a list of sentence candidates. Specifically, an NLP library (e.g., spaCy or GPT model) analyzes the word combinations and generates candidates as meaningful sentences.
[0177] Step 5:
[0178] The terminal presents the user with sentence suggestions sent from the server. In this step, the input is a list of generated sentence suggestions, and the output is the sentence displayed on the user's screen. Specifically, the terminal displays the sentence suggestions on its screen so that the user can visually confirm them.
[0179] Step 6:
[0180] The user selects a desired sentence from several sentence options displayed on the screen. In this step, the input is the sentence options displayed on the screen, and the output is the sentence selected by the user. Specifically, the user uses a touchscreen or mouse to select the desired sentence.
[0181] Step 7:
[0182] The terminal sends the text selected by the user to the server, which then distributes it to the robot control unit. In this step, the input is the text selected by the user, and the output is the control instruction for the robot. Specifically, the terminal sends the selected text to the server, and the server relays that instruction to the robot.
[0183] Step 8:
[0184] The robot performs specific operations based on the instructions it receives. In this step, the input is the instruction sent to the robot, and the output is the robot's physical movement. For example, based on the instruction "pick up and move the part," the robot arm picks up the part and moves it to the designated location.
[0185] This system enables intuitive robot operation using sign language, leading to increased efficiency in factory operations.
[0186] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0187] This invention combines a system that captures a user's sign language movements, analyzes the movement data to recognize specific words, with an emotion engine that recognizes the user's emotions. This system enables users to efficiently and intuitively generate sentences using sign language, and in addition, to generate sentences that include appropriate emotional expressions.
[0188] The user performs sign language actions through a device equipped with a camera and sensors. The device uses the camera to capture these sign language actions in real time. The camera continuously acquires video data and transmits it to a server.
[0189] The server analyzes the received video data and converts sign language movements into specific words. This uses image processing techniques and machine learning algorithms. Specifically, the server recognizes the shape and movement patterns of the hands and identifies the corresponding words based on a pre-trained model.
[0190] Next, the server generates sentence candidates based on the recognized words. Natural language processing (NLP) techniques are used here. The server generates multiple sentence candidates based on the recognized word combinations and lists them.
[0191] Furthermore, this system incorporates an emotion engine. The emotion engine analyzes the user's facial expressions and tone of voice to identify the user's emotional state. For example, if a user is smiling while performing sign language, the emotion engine recognizes that the user is "happy."
[0192] Based on the emotional state identified by the emotion engine, the server adjusts the tone and content of the generated text suggestions. For example, if the user is "happy," the server prioritizes generating text suggestions with a positive tone. Conversely, if the user is "sad," the server generates text suggestions that are more empathetic to that emotion.
[0193] The generated sentence candidates are sent to the terminal, which then visually presents them to the user. Specifically, the candidate sentences are displayed on the screen, allowing the user to select one. The user uses finger movements or touch gestures to select the desired sentence from the presented candidates.
[0194] The device detects the user's selection and displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others. The introduction of an emotion engine is expected to result in a more accurate and richer expression of the message.
[0195] For example, if a user performs the sign language actions for "I," "go," and "school," and the emotion engine recognizes the user's emotional state as "happy," the system will generate positive sentence options such as "I'm looking forward to going to school" or "I can't wait to go to school." The user then selects their desired sentence from these options, and the device displays the selected sentence as the final output.
[0196] In this way, the present invention is a system that can streamline communication using sign language and reflect the emotions of sign language users, thereby realizing more natural and enriching communication.
[0197] The following describes the processing flow.
[0198] Step 1:
[0199] The user performs sign language actions. A device equipped with a camera and sensors captures these sign language actions in real time. Specifically, the device uses its camera to continuously acquire video data, accurately capturing the shape and movement of the hand.
[0200] Step 2:
[0201] The terminal sends the acquired video data to the server. The server receives this video data and begins analysis. Data compression technology is used during data transfer for efficient transmission.
[0202] Step 3:
[0203] The server analyzes video data and converts sign language movements into specific words. The server uses image processing technology to analyze the contours and movements of the hands and maps the recognized movements to corresponding words through machine learning algorithms. Specifically, the server analyzes the shape, position, and movement patterns of the hands to recognize words such as "subject," "predicate," and "object."
[0204] Step 4:
[0205] The server uses natural language processing (NLP) techniques to generate sentence candidates based on the recognized words. For example, from the recognized words "I," "go," and "school," several sentences such as "I go to school" and "It is I who go to school" are created. The server then lists these sentence candidates.
[0206] Step 5:
[0207] The emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotions. The device uses its camera and microphone to capture the user's facial expressions and voice data, and sends this data to the server.
[0208] Step 6:
[0209] The server analyzes the facial and voice data it receives to identify the user's emotional state. The emotion engine uses facial recognition and voice analysis technologies to recognize the user's emotional state, such as "happy," "sad," or "surprised."
[0210] Step 7:
[0211] The server adjusts the tone and content of the generated sentence suggestions based on the identified emotional state. For example, if the user is perceived as "happy," the server prioritizes generating sentence suggestions with a positive tone. Specifically, it might generate sentence suggestions that include emotional expressions such as, "I'm looking forward to going to school."
[0212] Step 8:
[0213] The server sends generated sentiment-reflecting sentence candidates to the terminal, which then visually presents them to the user. Multiple candidate sentences are displayed on the screen in a list format, with each candidate numbered or highlighted.
[0214] Step 9:
[0215] The user selects their desired text from a list of suggested texts. The user makes the selection using finger movements or touch gestures. The device captures the user's selection actions and sends that information to the server.
[0216] Step 10:
[0217] The server receives the user's selection information and finally confirms the selected text. The selected text is then sent from the server to the terminal.
[0218] Step 11:
[0219] The device visually displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others.
[0220] Through the processing steps described above, the present invention realizes efficient and intuitive text generation and communication using sign language. Furthermore, the introduction of an emotion engine enables natural and rich communication that also reflects emotional expression.
[0221] (Example 2)
[0222] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0223] Conventional sign language recognition systems only recognize sign language movements, and the generated text does not reflect the user's emotions, resulting in the problem of inaccurately conveying the intended message. Furthermore, providing a diverse range of example sentences during text generation is difficult, limiting communication possibilities.
[0224] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words, means for presenting the generated sentence candidates to the user, means for the user to select a generated sentence candidate, and means for recognizing the user's emotional state and adjusting the tone and content of the sentence candidates based on that emotional state. This makes it possible to generate accurate sentences based on the user's sign language actions and to provide a variety of sentence candidates that reflect the user's emotions.
[0225] A "user" is a person who uses this system and is the entity that inputs data through sign language actions.
[0226] "Means for capturing movements" refers to devices or technologies that use cameras or sensors to acquire a user's sign language movements in real time.
[0227] "Motion data" refers to video information and sensor data that digitally records the user's sign language movements.
[0228] "Means of analysis" refers to methods and techniques for analyzing acquired motion data using machine learning algorithms and image processing technologies, and converting it into specific words.
[0229] A "specific word" is a string of characters with meaning recognized from the analyzed motion data, representing a conversion of sign language movements into language.
[0230] "Methods for generating sentence candidates" refer to methods and techniques for creating multiple sentences based on recognized words using natural language processing technology or generative AI models.
[0231] "Means for presenting generated sentence candidates to the user" refers to a technology that visually displays multiple sentence candidates to the user using a display device such as a screen.
[0232] "Means for users to select generated text candidates" refers to a device or technology that allows a user to select a desired text from a visually displayed list of text candidates using a touchscreen or finger movements.
[0233] "Means of recognizing emotional states" refers to methods and technologies that use facial recognition technology or voice analysis technology to identify a user's emotions.
[0234] "Means of adjusting tone and content" refers to methods and techniques for appropriately modifying the style and content of candidate texts generated based on the recognized emotional state of the user.
[0235] This invention relates to a system that captures a user's sign language movements, analyzes the movement data to recognize specific words, and further recognizes the user's emotions to generate sentence candidates based on those emotions. This system efficiently and intuitively generates sentences based on the user's sign language movements and reflects emotions, enabling natural communication.
[0236] hardware
[0237] The system primarily uses the following hardware:
[0238] 1. Device: An electronic device equipped with a camera or sensors (e.g., smartphone, tablet, etc.).
[0239] 2. Camera: A high-resolution camera that captures sign language movements in real time.
[0240] software
[0241] The system uses the following software and technologies:
[0242] 1. Image processing technology: Use OpenCV or similar tools to analyze sign language movements.
[0243] 2. Machine learning algorithm: TensorFlow or PyTorch is used to recognize specific words from behavioral data.
[0244] 3. Natural Language Processing (NLP) Technology: Generative AI models such as GPT-3 (registered trademark) are used to generate sentence candidates based on recognized words.
[0245] 4. Emotion Recognition Technology: Use Amazon Rekognition or Azure® Face API to identify the user's emotional state and adjust suggested text accordingly.
[0246] Data processing and calculation
[0247] Capture and transmit sign language movements
[0248] The user performs sign language gestures in front of the device. The device uses its built-in camera to capture the sign language gestures in real time and sends the video data frame by frame to the server. HTTP or WebSocket protocol is used for transmission.
[0249] Video data analysis and word recognition
[0250] The server analyzes the received video data using image processing technologies such as OpenCV and TensorFlow, as well as machine learning algorithms, to recognize specific words. It detects hand shapes and movement patterns and converts them into corresponding words based on a pre-trained model.
[0251] Sentence candidate generation
[0252] The server generates sentence candidates using a generative AI model such as GPT-3 based on recognized words. It takes a prompt message as input and generates and lists multiple sentence candidates. For example, if the user inputs the words "I," "go," and "school," the server will use a prompt message like the following:
[0253] Sign language words to input: I, go, school
[0254] Emotional state: Happy
[0255] Generation example:
[0256] I look forward to going to school.
[0257] I can't wait to go to school.
[0258] Identifying and adjusting emotional states
[0259] The server uses emotion recognition technologies such as Amazon Rekognition and Azure Face API to analyze the user's facial expressions and voice tone from video data and identify their emotional state. For example, if a user is smiling while using sign language, the server recognizes this as "happy." Based on the identified emotional state, the server adjusts the tone and content of the generated text suggestions.
[0260] Presentation and selection of sentence options
[0261] The generated sentence suggestions are sent to the device, which displays them on its screen. The user selects the desired sentence using touch gestures or finger movements. For example, "I'm looking forward to going to school" and "I can't wait to go to school" might be displayed, and the user selects from among them.
[0262] Display of final output
[0263] The device ultimately displays the text selected by the user. This allows the user to accurately convey what they want to communicate, including their emotions, to others.
[0264] As described above, the system of the present invention enables natural communication that reflects the user's sign language movements and emotions.
[0265] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0266] Step 1: The user performs a sign language action.
[0267] The user performs sign language actions in front of the device's camera. For example, they might perform sign language actions for "I," "go," and "school." The user's sign language actions are the input and are captured as video data.
[0268] Step 2: The device captures sign language movements and sends the video data to the server.
[0269] The device uses a built-in high-resolution camera to capture the user's sign language movements in real time. The captured video data is acquired frame by frame and transmitted to the server via Wi-Fi or mobile data communication. The input is video data of the sign language movements, and the output is video data transferred to the server.
[0270] Step 3: The server analyzes the video data and converts sign language actions into words.
[0271] The server uses image processing techniques such as OpenCV and TensorFlow / PyTorch, along with machine learning algorithms, to analyze the received video data. It detects hand shapes and movement patterns and uses a pre-trained model to convert those movements into corresponding specific words. The input is video data, and the output is a list of recognized words.
[0272] Step 4: The server generates sentence suggestions based on the words.
[0273] The server generates sentence candidates using a generative AI model (e.g., GPT-3) based on recognized word combinations. A prompt sentence is entered, and multiple sentence candidates are generated and listed. The input consists of a list of recognized words and a prompt sentence, and the output is a list of sentence candidates. For example, if the words "I," "go," and "school" are entered, the server generates sentence candidates such as "I look forward to going to school" and "I can't wait to go to school."
[0274] Step 5: The server uses the emotion engine to identify the emotional state and adjust the suggested sentences.
[0275] The server uses emotion recognition technology (such as Amazon Rekognition or Azure Face API) to analyze facial expressions and voice tone to identify the user's emotional state. For example, if the user is smiling and using sign language in video data, it will recognize that the user is "happy." The input is the user's video data, and the output is the identified emotional state. Subsequently, the tone and content of the generated sentence candidates are adjusted according to the emotional state. For example, if the user is happy, positive sentences will be prioritized.
[0276] Step 6: The server sends the adjusted text candidates to the terminal.
[0277] The server sends the adjusted sentence candidates to the terminal in JSON format. The input is the adjusted sentence candidates, and the output is the data sent to the terminal. Specifically, sentence candidates such as "I'm looking forward to going to school" and "I can't wait to go to school" are sent.
[0278] Step 7: The device presents text suggestions to the user, and the user makes a selection.
[0279] The terminal displays the received sentence candidates on the display. The user selects the desired sentence using the touch screen or finger movement. Sentence candidates are sent to the terminal as input, and the sentence selected by the user is obtained as output. For example, multiple sentences are displayed on the display, and an operation where the user selects "I am looking forward to going to school" from them is included.
[0280] Step 8: The terminal displays the final sentence
[0281] The terminal finally displays the sentence selected by the user on the display. The sentence selected by the user is the input, and the final displayed sentence is obtained as output. For example, the sentence "I am looking forward to going to school" is displayed on the display. With this display, the user can accurately convey the content they want to convey to others.
[0282] (Application Example 2)
[0283] Next, Application Example 2 will be described. In the following description, the data processing device 12 is referred to as a "server", and the smart device 14 is referred to as a "terminal".
[0284] The present invention aims to improve the efficiency and accuracy of communication at the work site by real-time textifying sign language and generating text taking into account the emotional state of the operator when a worker with hearing impairment communicates using sign language. Also, as an issue, it is to reduce the stress and anxiety of the worker by providing appropriate feedback and support reflecting the emotional state, and to improve the safety and productivity of the work.
[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following respective means.
[0286] In this invention, the server includes means for capturing the user's actions, means for analyzing the captured action data to recognize specific words, means for generating sentence candidates based on the recognized words and the user's emotional state, means for presenting the generated sentence candidates to the user, and means for the user to select the generated sentence candidates. As a result, communication using sign language is made more efficient and the user's emotional state is also reflected, enabling richer and more accurate communication.
[0287] The "means for capturing the user's actions" is a device that uses a camera or sensor to acquire the user's physical actions as digital data.
[0288] The "means for analyzing the captured action data to recognize specific words" is a device or software that uses machine learning algorithms or image processing techniques to identify specific words or meanings based on the user's action data.
[0289] The "means for generating sentence candidates based on the recognized words and the user's emotional state" is a device or software that uses the words recognized from the user's sign language actions and the user's emotional state analyzed by an emotion engine to generate multiple sentence candidates.
[0290] The "means for presenting the generated sentence candidates to the user" is a display that visually presents options of the generated sentences to the user or a device that provides voice notifications.
[0291] The "means for the user to select the generated sentence candidates" is an interface for the user to select a desired sentence from the presented sentence candidates, referring to a touch screen or other input devices.
[0292] "Sign language actions" are a means of communication that uses movements of the hands and fingers and are physical actions for conveying specific words or meanings.
[0293] "Emotional state" refers to the user's emotional and mood state, and is identified by analyzing facial expressions, tone of voice, and other factors.
[0294] "Natural language processing technology" refers to technologies that enable computers to understand and generate human language, and includes algorithms for text analysis and generation.
[0295] "Emotional analysis technology" is a technology that analyzes a user's facial expressions, voice tone, body movements, etc., to identify the user's current emotional state.
[0296] The embodiments for carrying out this invention will be described in detail below.
[0297] First, the system program is generated. This program captures the user's sign language movements, analyzes the movement data, and recognizes specific words. It also incorporates an emotion engine that recognizes the user's emotional state, and includes emotional expressions in the generated sentences.
[0298] The system consists of the following main hardware and software components:
[0299] 1. Means for capturing user actions
[0300] Cameras and sensors are used to capture the user's hand and finger movements in real time. Specifically, webcams and dedicated motion capture devices are used.
[0301] 2. A means of analyzing captured motion data and recognizing specific words.
[0302] Machine learning algorithms and image processing techniques are used to analyze captured sign language movements. Specifically, libraries such as TensorFlow and OpenCV are used.
[0303] 3. A means of generating sentence candidates based on recognized words and the user's emotional state.
[0304] Generate a plurality of candidate sentences using natural language processing (NLP) technology and sentiment analysis technology. Specifically, use an emotion engine (e.g., EmotionEngine) and an NLP model (e.g., NLPModel).
[0305] 4. Means for presenting the generated candidate sentences to the user
[0306] Use a display to present the generated candidate sentences to the user. For this, a computer monitor, a projector, etc. are used.
[0307] 5. Means for the user to select the generated candidate sentences
[0308] Use a touch screen or other interface for the user to select the desired sentence. Specifically, a touch panel or a pointing device is used.
[0309] Furthermore, specific examples are shown below.
[0310] As an example, consider the case where a worker with hearing impairment working in a factory conveys "The machine is malfunctioning" in sign language. First, a camera captures the sign language motion, and the system analyzes the motion and recognizes "The machine is malfunctioning". At the same time, the emotion engine analyzes the worker's expression and reads an anxious emotion. Based on this, the system generates a candidate sentence "The machine is malfunctioning. I'm very anxious" and displays it on the display. The worker can select this sentence using the touch panel and inform other workers or managers.
[0311] Examples of prompt sentences are as follows:
[0312] "Imagine a system that uses sign language to report the situation within a factory. For example, if a worker expresses in sign language that 'the machine is broken' and shows an anxious expression, build a system that analyzes the sign language and emotions to provide appropriate feedback. This system combines sign language recognition and emotion recognition technologies to generate a detailed report."
[0313] Thus, this invention streamlines communication using sign language and reflects the emotions of the workers, thereby achieving more natural and enriching communication.
[0314] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0315] Step 1:
[0316] Camera-based motion capture
[0317] The server uses a camera to capture the user's sign language movements in real time. The input is the user's hand and finger movements, and the output is video frame data. This frame data is then sent to the next analysis step.
[0318] Step 2:
[0319] Sign language motion recognition
[0320] The server analyzes captured video frame data using image processing techniques and machine learning algorithms to recognize specific words. At this stage, TensorFlow and OpenCV are used to analyze hand shapes and movement patterns. The input is video frame data, and the output is the recognized word.
[0321] Step 3:
[0322] Emotion analysis
[0323] The server analyzes the user's facial expressions and voice tone from the captured video frame data and uses an EmotionEngine to identify the user's emotional state. The input is the same video frame data, and the output is the analyzed emotional state.
[0324] Step 4:
[0325] Sentence generation
[0326] The server generates multiple sentence candidates using natural language processing (NLP) techniques based on recognized words and analyzed sentiment states. An NLP model is used at this stage. The input is recognized words and sentiment states, and the output is multiple sentence candidates.
[0327] Step 5:
[0328] Suggestion of sentence candidates
[0329] The terminal visually displays the generated sentence candidates on its screen and presents them to the user. The input consists of multiple sentence candidates, and the output includes the sentence candidates displayed on the screen. The user is then ready to make a selection.
[0330] Step 6:
[0331] Selection of sentence candidates
[0332] The user selects their desired sentence from the presented options. The device acquires the user's selection data using a touchscreen or pointing device. The input is the user's selection operation, and the output is the final sentence selected by the user.
[0333] Step 7:
[0334] Output of selected text
[0335] The device displays the final text selected by the user on its screen, confirming the intent of the communication. The input is the text selected by the user, and the output is the text that is finally displayed on the device. This ensures that the user's intent is conveyed to the other person.
[0336] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0337] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0338] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0339] [Second Embodiment]
[0340] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0341] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0342] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0343] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0344] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0345] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0346] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0347] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0348] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0349] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0350] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0351] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0352] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate and communicate text using sign language.
[0353] Users perform sign language actions through a device equipped with a camera and sensors. The device uses the camera to capture these actions in real time. The camera continuously acquires video data and transmits it to a server.
[0354] The server analyzes the received video data and converts sign language movements into specific words. This uses image processing techniques and machine learning algorithms. Specifically, the server recognizes the shape and movement patterns of the hands and identifies the corresponding words based on a pre-trained model.
[0355] Next, the server generates sentence candidates based on the recognized words. Natural language processing (NLP) techniques are used here. The server generates multiple sentence candidates based on the recognized word combinations and lists them.
[0356] The generated sentence candidates are sent to the terminal, which then visually presents them to the user. Specifically, the candidate sentences are displayed on the screen, allowing the user to select one. The user uses finger movements or touch gestures to select the desired sentence from the presented candidates.
[0357] The device detects the user's selection and displays the final selected sentence. This allows the user to easily confirm the message they want to convey.
[0358] For example, if a user performs the sign language actions for "I," "go," and "school," the server recognizes these as specific words and generates several sentence options, such as "I go to school," "It is I who goes to school," and "The place I go to is school." The user then selects their desired sentence from these options, and the terminal displays the selected sentence as the final output.
[0359] In this way, the present invention is a system that can streamline communication using sign language and significantly reduce the burden on sign language users.
[0360] The following describes the processing flow.
[0361] Step 1:
[0362] The user performs sign language actions. A device equipped with a camera and sensors captures these sign language actions in real time. Specifically, the device continuously acquires video data, accurately capturing the shape and movement of the hand.
[0363] Step 2:
[0364] The terminal sends the acquired video data to the server. The server receives this video data and begins analysis. At this stage, data preprocessing such as noise reduction and color correction may be performed.
[0365] Step 3:
[0366] The server analyzes video data and converts sign language movements into specific words. The server uses image processing technology to analyze the contours and movements of the hands and maps the recognized movements to corresponding words through machine learning algorithms. This identifies each element of the sign language (e.g., subject, predicate, object, etc.).
[0367] Step 4:
[0368] The server uses natural language processing (NLP) techniques to generate sentence candidates based on the words it recognizes. For example, from the words "I," "go," and "school," multiple sentences such as "I go to school" and "It is I who go to school" are created.
[0369] Step 5:
[0370] The server sends generated sentence candidates to the terminal, which then visually presents them to the user. Specifically, multiple candidate sentences are displayed on the screen in a list format. Each sentence candidate is identified using a number or highlighting.
[0371] Step 6:
[0372] The user selects their desired text from a list of suggested texts. The user makes the selection using finger movements or touch gestures. The device captures the user's selection actions and sends that information to the server.
[0373] Step 7:
[0374] The server receives the user's selection information and finally confirms the selected text. The selected text is then sent from the server to the terminal.
[0375] Step 8:
[0376] The device visually displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others.
[0377] Through the processing steps described above, the present invention realizes efficient and intuitive text generation and communication using sign language.
[0378] (Example 1)
[0379] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0380] While modern communication largely relies on text-based technologies, their use can be limited for users who primarily rely on sign language, such as those with hearing impairments. Therefore, there is a need for systems that can efficiently and intuitively generate text using sign language and transmit it directly to others. However, existing technologies have been insufficient in accurately recognizing and transcribing sign language movements, hindering smooth communication for sign language users. In particular, challenges existed in the accurate capture and analysis of sign language movements, and the subsequent natural text generation process.
[0381] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0382] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, and means for generating sentence candidates based on the recognized words. This enables the real-time capture of the user's sign language actions, transmission of the action data to the server, and analysis of the sign language actions using image processing technology and machine learning algorithms. Furthermore, by utilizing natural language processing technology to generate multiple sentence candidates and presenting them visually to the user, the user can easily select their desired sentence and confirm the final sentence. This enables sign language users, including those with hearing impairments, to communicate more efficiently and intuitively.
[0383] "Means for capturing user actions" refer to devices or methods for recording a user's sign language movements in real time, and generally involve the use of cameras or sensors.
[0384] "Means for analyzing captured motion data and recognizing specific words" refers to technologies for analyzing acquired video data of sign language movements and identifying specific words based on that analysis, and includes image processing technologies and machine learning algorithms.
[0385] "A means of generating sentence candidates based on recognized words" refers to a system that automatically generates appropriate sentences by combining words recognized through analysis, and utilizes natural language processing technology.
[0386] "Means for presenting generated text candidates to the user" refers to devices or methods for displaying text candidates in a way that the user can easily review, and primarily uses displays or monitors.
[0387] "Means for the user to select generated sentence candidates" refers to an interface for the user to choose their preferred sentence from among several presented candidates, and includes touchscreens, pointing devices, and the like.
[0388] "Means for displaying selected text" refers to devices or methods for visually showing the text ultimately selected by the user, and displays or screens are used for this purpose.
[0389] "Methods of using cameras and sensors for capture" refers to technologies that utilize cameras and various sensors to record detailed sign language movements.
[0390] "Image processing techniques for identifying hand shape and movement patterns" are methods for characterizing hand shape and movement by analyzing data acquired from cameras and sensors, and utilize image processing libraries, etc.
[0391] "Methods for recognizing words using trained machine learning algorithms" refers to technologies that automatically identify corresponding words from sign language gestures using machine learning models that have been trained on a large amount of data in advance.
[0392] "Methods for generating sentence candidates using natural language processing technology" refers to methods that utilize machine learning models to generate multiple natural-sounding sentences from recognized word combinations.
[0393] "Means for a terminal to send video data acquired by its camera to a server" refers to a technology that allows a terminal to send video data captured in real time to a server via communication means such as the internet.
[0394] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate and communicate text using sign language. Specific embodiments are described below.
[0395] Hardware and software configuration
[0396] Hardware:
[0397] 1. Camera: Used to capture the user's sign language movements. Generally, a high-resolution camera is used.
[0398] 2. Sensors: Used to help accurately determine the position of hands and fingers. Depth sensors are examples of such sensors.
[0399] 3. Terminal: A device used to process captured video data and display the results to the user. Smartphones and tablets are typical examples.
[0400] 4. Display: Built into the device and used to visually present text suggestions to the user.
[0401] software:
[0402] 1. Image processing technology: Using libraries such as OpenCV, we extract the shape and movement characteristics of the hand from video data captured by the camera.
[0403] 2. Machine Learning Algorithms: Using TensorFlow, PyTorch, etc., a model is constructed and applied to convert sign language actions into specific words based on the extracted features.
[0404] 3. Natural Language Processing (NLP): Using generative AI models such as GPT and BERT, natural-sounding sentence candidates are generated based on recognized words.
[0405] Specific description of the system's operation
[0406] The user stands in front of the device's camera to perform sign language actions. The device's camera captures the user's sign language actions in real time. The video data acquired by the camera is transmitted to a server via the internet. The server analyzes the received video data and uses image processing technology to identify hand shapes and movement patterns. Next, a machine learning algorithm is applied to convert the identified sign language actions into corresponding words.
[0407] Next, natural language processing technology is used to generate multiple sentence candidates based on the recognized words. The generated sentence candidates are sent from the server to the terminal and displayed on the terminal's screen. The user selects the desired sentence from the presented candidates using touch gestures or other methods. Finally, the selected sentence is displayed on the terminal's screen, allowing the user to show the content they want to communicate to others.
[0408] Specific example
[0409] As an example, consider a case where a user performs the sign language actions for "I," "go," and "school." The device's camera captures these actions and sends the data to a server. The server analyzes the video data and identifies the words "I," "go," and "school." Next, using natural language processing technology, it generates several sentence options, such as "I go to school," "It is I who goes to school," and "The place I go to is school." The generated sentence options are sent to the device and displayed on the screen. The user selects their desired sentence, and the final sentence is displayed on the screen.
[0410] Example of a prompt
[0411] Example: "Please write a prompt for a system that analyzes sign language movements and generates multiple sentence options based on them."
[0412] As described above, the system of the present invention can significantly improve communication using sign language and reduce the communication burden on sign language users.
[0413] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0414] Step 1: Capture the user's sign language movements
[0415] Input: User's sign language gestures
[0416] Specific operation: The user stands in front of the device's camera to perform sign language actions. The device's camera captures the sign language actions in real time and continuously acquires video data.
[0417] Output: Acquired video data
[0418] Step 2: Sending video data
[0419] Input: Acquired video data
[0420] Specific operation: The device sends the captured video data to the server. The data is streamed to the server in real time over the internet.
[0421] Output: Video data sent to the server
[0422] Step 3: Analyzing video data
[0423] Input: Video data sent to the server
[0424] Specific operation: The server analyzes the received video data. First, it uses image processing techniques (such as OpenCV) to extract features of the hand's shape and movement. Afterwards, it identifies feature quantities (hand shape, position, movement pattern, etc.).
[0425] Output: Feature data
[0426] Step 4: Word Recognition
[0427] Input: Feature data
[0428] Specific operation: The server applies machine learning algorithms (TensorFlow or PyTorch) based on the extracted features to convert sign language actions into corresponding words. A pre-trained model identifies the appropriate word from the sign language action.
[0429] Output: Recognized words
[0430] Step 5: Generating text candidates
[0431] Input: Recognized word
[0432] Specific operation: The server generates multiple sentence candidates using natural language processing (NLP) techniques (GPT and BERT) based on recognized words. It then automatically generates appropriate sentence combinations using a generative AI model.
[0433] Output: Sentence suggestion list
[0434] Step 6: Send the suggested message
[0435] Input: Sentence suggestion list
[0436] Specific operation: The server sends a list of generated sentence suggestions to the terminal. The data is sent to the terminal via the network.
[0437] Output: List of suggested sentences sent to the terminal
[0438] Step 7: Display candidate sentences
[0439] Input: List of suggested sentences sent to the terminal
[0440] Specific operation: The device displays received text suggestions on its screen. It displays a list so that the user can visually see multiple text suggestions.
[0441] Output: Suggested sentences displayed on the screen
[0442] Step 8: User Selection
[0443] Input: Suggested sentences displayed on the screen
[0444] Specific operation: The user selects the desired sentence from the presented sentence options using finger movements or touch operations.
[0445] Output: Selected text
[0446] Step 9: Display the final sentence
[0447] Input: Selected text
[0448] Specific operation: The device displays the final sentence selected by the user on its screen. This allows the user to show others the message they ultimately wanted to convey.
[0449] Output: The final sentence displayed on the screen
[0450] (Application Example 1)
[0451] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0452] Traditional factory robot operation typically involves keyboards or touch panels, limiting intuitive control methods. This posed a problem, particularly for workers whose primary means of communication is sign language, as the operation was complex and difficult to understand. Furthermore, the lack of sufficient means to improve work efficiency was also a challenge.
[0453] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0454] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words, means for presenting the generated sentence candidates to the user, means for the user to select a generated sentence candidate, means for the user to give instructions for robot operation using sign language, and means for robot operation to execute the recognized sign language instructions. This makes it possible to operate the robot intuitively and efficiently using sign language.
[0455] "Means for capturing user movements" refers to devices that use cameras or sensors to capture the user's body movements and acquire those movements as digital data.
[0456] "A means of analyzing captured motion data and recognizing specific words" refers to the process of analyzing acquired digital data and identifying words that correspond to specific sign language or actions intended by the user.
[0457] "Methods for generating sentence candidates based on recognized words" refers to algorithms that combine words identified through analysis to generate candidates that form natural-sounding sentences.
[0458] "Means of presenting generated sentence candidates to the user" refers to an interface that visually displays multiple generated sentences as candidates on a display or screen, allowing the user to select one.
[0459] "Means for the user to select generated sentence candidates" refers to the means by which the user selects the desired sentence from among multiple displayed sentence candidates.
[0460] "A means by which a user can give instructions to operate a robot using sign language" refers to an interface for transmitting operating instructions to a robot or machine through sign language actions.
[0461] A "robot operating means for executing instructions based on recognized sign language" is a control device that allows a robot to perform specific operations based on instructions recognized as sign language actions.
[0462] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate text using sign language and to operate robots and give instructions in a factory setting.
[0463] hardware
[0464] The server and terminal include the following hardware:
[0465] Camera (e.g., Logitech C920): Captures the user's sign language movements in real time.
[0466] Robot control device: A device for controlling a robot based on recognized sign language movements.
[0467] software
[0468] The following software will be used:
[0469] Python: A programming language used for the overall system implementation.
[0470] OpenCV (CV2): A library for capturing and pre-processing camera footage.
[0471] Keras: A machine learning framework for running sign language gesture recognition models.
[0472] Natural language processing libraries (e.g., spaCy, GPT model): These are libraries for generating text candidates based on sign language actions.
[0473] System operation
[0474] The server receives video data of sign language movements transmitted from the user's camera. The video data is first preprocessed using OpenCV and then fed into a Keras model. The Keras model uses a pre-trained machine learning algorithm to convert the sign language movements into corresponding words.
[0475] The converted words are then transformed into multiple sentence candidates using a natural language processing library (e.g., spaCy or the GPT model). These sentence candidates are sent from the server to the terminal and displayed visually on the terminal's screen.
[0476] The user selects their desired sentence from several sentence options displayed on the screen. This selection is made through touch operations or finger movements. The final selected sentence is displayed on the device and then sent to the robot.
[0477] The robot performs specific actions based on the selected sentence. For example, if the user performs the sign language action for "pick up the part and move it," the robot will receive instructions such as "pick up the part and move it 5 meters."
[0478] Specific example
[0479] When a worker performs a sign language gesture for "pick up the part and move it," the server analyzes it and recognizes it as a corresponding word. The server then uses natural language processing technology to generate sentence options such as "pick up the part and move it 5 meters" or "take the part and move it to the designated location." These options are sent to the terminal, and the user selects the one sentence they prefer. Finally, the robot performs the action based on the selected instruction.
[0480] Example of a prompt
[0481] Example of sign language recognition results for use in analyzing sign language movements and generating text: Please generate phrases corresponding to the sign language words "part," "pick up," and "move." Possible instruction sentences would be something like, "Pick up the part and move 5 meters."
[0482] This invention enables intuitive and efficient operation and instruction using sign language in factories. Furthermore, it makes operation easier and improves work efficiency for workers who primarily use sign language for communication.
[0483] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0484] Step 1:
[0485] The server receives video data of sign language movements transmitted from the user's camera in real time. This video data is data captured by the camera of the sign language movements performed by the user. The data input is a video stream from the camera, and the output is the frame data of that video. Specifically, the camera continuously captures the user's hand movements and sends each frame to the server as digital data.
[0486] Step 2:
[0487] The server analyzes the received video data. OpenCV is used for preprocessing the video. In this step, the necessary sign language movements are extracted from the video, and noise is removed. The input for this process is frame data, and the output is preprocessed image data of the sign language movements. Specifically, this involves image smoothing and edge detection to enhance the shape of the hands.
[0488] Step 3:
[0489] The server supplies pre-processed video data to a Keras model, which runs a machine learning algorithm that converts sign language gestures into corresponding words. In this step, the input is pre-processed image data, and the output is the recognized words. Specifically, the model analyzes certain features in the image and maps them to pre-trained words.
[0490] Step 4:
[0491] The server generates multiple sentence candidates using a natural language processing library based on the recognized words. In this step, the input is the recognized words, and the output is a list of sentence candidates. Specifically, an NLP library (e.g., spaCy or GPT model) analyzes the word combinations and generates candidates as meaningful sentences.
[0492] Step 5:
[0493] The terminal presents the user with sentence suggestions sent from the server. In this step, the input is a list of generated sentence suggestions, and the output is the sentence displayed on the user's screen. Specifically, the terminal displays the sentence suggestions on its screen so that the user can visually confirm them.
[0494] Step 6:
[0495] The user selects a desired sentence from several sentence options displayed on the screen. In this step, the input is the sentence options displayed on the screen, and the output is the sentence selected by the user. Specifically, the user uses a touchscreen or mouse to select the desired sentence.
[0496] Step 7:
[0497] The terminal sends the text selected by the user to the server, which then distributes it to the robot control unit. In this step, the input is the text selected by the user, and the output is the control instruction for the robot. Specifically, the terminal sends the selected text to the server, and the server relays that instruction to the robot.
[0498] Step 8:
[0499] The robot performs specific operations based on the instructions it receives. In this step, the input is the instruction sent to the robot, and the output is the robot's physical movement. For example, based on the instruction "pick up and move the part," the robot arm picks up the part and moves it to the designated location.
[0500] This system enables intuitive robot operation using sign language, leading to increased efficiency in factory operations.
[0501] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0502] This invention combines a system that captures a user's sign language movements, analyzes the movement data to recognize specific words, with an emotion engine that recognizes the user's emotions. This system enables users to efficiently and intuitively generate sentences using sign language, and in addition, to generate sentences that include appropriate emotional expressions.
[0503] The user performs sign language actions through a device equipped with a camera and sensors. The device uses the camera to capture these sign language actions in real time. The camera continuously acquires video data and transmits it to a server.
[0504] The server analyzes the received video data and converts sign language movements into specific words. This uses image processing techniques and machine learning algorithms. Specifically, the server recognizes the shape and movement patterns of the hands and identifies the corresponding words based on a pre-trained model.
[0505] Next, the server generates sentence candidates based on the recognized words. Natural language processing (NLP) techniques are used here. The server generates multiple sentence candidates based on the recognized word combinations and lists them.
[0506] Furthermore, this system incorporates an emotion engine. The emotion engine analyzes the user's facial expressions and tone of voice to identify the user's emotional state. For example, if a user is smiling while performing sign language, the emotion engine recognizes that the user is "happy."
[0507] Based on the emotional state identified by the emotion engine, the server adjusts the tone and content of the generated text suggestions. For example, if the user is "happy," the server prioritizes generating text suggestions with a positive tone. Conversely, if the user is "sad," the server generates text suggestions that are more empathetic to that emotion.
[0508] The generated sentence candidates are sent to the terminal, which then visually presents them to the user. Specifically, the candidate sentences are displayed on the screen, allowing the user to select one. The user uses finger movements or touch gestures to select the desired sentence from the presented candidates.
[0509] The device detects the user's selection and displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others. The introduction of an emotion engine is expected to result in a more accurate and richer expression of the message.
[0510] For example, if a user performs the sign language actions for "I," "go," and "school," and the emotion engine recognizes the user's emotional state as "happy," the system will generate positive sentence options such as "I'm looking forward to going to school" or "I can't wait to go to school." The user then selects their desired sentence from these options, and the device displays the selected sentence as the final output.
[0511] In this way, the present invention is a system that can streamline communication using sign language and reflect the emotions of sign language users, thereby realizing more natural and enriching communication.
[0512] The following describes the processing flow.
[0513] Step 1:
[0514] The user performs sign language actions. A device equipped with a camera and sensors captures these sign language actions in real time. Specifically, the device uses its camera to continuously acquire video data, accurately capturing the shape and movement of the hand.
[0515] Step 2:
[0516] The terminal sends the acquired video data to the server. The server receives this video data and begins analysis. Data compression technology is used during data transfer for efficient transmission.
[0517] Step 3:
[0518] The server analyzes video data and converts sign language movements into specific words. The server uses image processing technology to analyze the contours and movements of the hands and maps the recognized movements to corresponding words through machine learning algorithms. Specifically, the server analyzes the shape, position, and movement patterns of the hands to recognize words such as "subject," "predicate," and "object."
[0519] Step 4:
[0520] The server uses natural language processing (NLP) techniques to generate sentence candidates based on the recognized words. For example, from the recognized words "I," "go," and "school," several sentences such as "I go to school" and "It is I who go to school" are created. The server then lists these sentence candidates.
[0521] Step 5:
[0522] The emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotions. The device uses its camera and microphone to capture the user's facial expressions and voice data, and sends this data to the server.
[0523] Step 6:
[0524] The server analyzes the facial and voice data it receives to identify the user's emotional state. The emotion engine uses facial recognition and voice analysis technologies to recognize the user's emotional state, such as "happy," "sad," or "surprised."
[0525] Step 7:
[0526] The server adjusts the tone and content of the generated sentence suggestions based on the identified emotional state. For example, if the user is perceived as "happy," the server prioritizes generating sentence suggestions with a positive tone. Specifically, it might generate sentence suggestions that include emotional expressions such as, "I'm looking forward to going to school."
[0527] Step 8:
[0528] The server sends generated sentiment-reflecting sentence candidates to the terminal, which then visually presents them to the user. Multiple candidate sentences are displayed on the screen in a list format, with each candidate numbered or highlighted.
[0529] Step 9:
[0530] The user selects their desired text from a list of suggested texts. The user makes the selection using finger movements or touch gestures. The device captures the user's selection actions and sends that information to the server.
[0531] Step 10:
[0532] The server receives the user's selection information and finally confirms the selected text. The selected text is then sent from the server to the terminal.
[0533] Step 11:
[0534] The device visually displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others.
[0535] Through the processing steps described above, the present invention realizes efficient and intuitive text generation and communication using sign language. Furthermore, the introduction of an emotion engine enables natural and rich communication that also reflects emotional expression.
[0536] (Example 2)
[0537] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0538] Conventional sign language recognition systems only recognize sign language movements, and the generated text does not reflect the user's emotions, resulting in the problem of inaccurately conveying the intended message. Furthermore, providing a diverse range of example sentences during text generation is difficult, limiting communication possibilities.
[0539] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words, means for presenting the generated sentence candidates to the user, means for the user to select a generated sentence candidate, and means for recognizing the user's emotional state and adjusting the tone and content of the sentence candidates based on that emotional state. This makes it possible to generate accurate sentences based on the user's sign language actions and to provide a variety of sentence candidates that reflect the user's emotions.
[0540] A "user" is a person who uses this system and is the entity that inputs data through sign language actions.
[0541] "Means for capturing movements" refers to devices or technologies that use cameras or sensors to acquire a user's sign language movements in real time.
[0542] "Motion data" refers to video information and sensor data that digitally records the user's sign language movements.
[0543] "Means of analysis" refers to methods and techniques for analyzing acquired motion data using machine learning algorithms and image processing technologies, and converting it into specific words.
[0544] A "specific word" is a string of characters with meaning recognized from the analyzed motion data, representing a conversion of sign language movements into language.
[0545] "Methods for generating sentence candidates" refer to methods and techniques for creating multiple sentences based on recognized words using natural language processing technology or generative AI models.
[0546] "Means for presenting generated sentence candidates to the user" refers to a technology that visually displays multiple sentence candidates to the user using a display device such as a screen.
[0547] "Means for users to select generated text candidates" refers to a device or technology that allows a user to select a desired text from a visually displayed list of text candidates using a touchscreen or finger movements.
[0548] "Means of recognizing emotional states" refers to methods and technologies that use facial recognition technology or voice analysis technology to identify a user's emotions.
[0549] "Means of adjusting tone and content" refers to methods and techniques for appropriately modifying the style and content of candidate texts generated based on the recognized emotional state of the user.
[0550] This invention relates to a system that captures a user's sign language movements, analyzes the movement data to recognize specific words, and further recognizes the user's emotions to generate sentence candidates based on those emotions. This system efficiently and intuitively generates sentences based on the user's sign language movements and reflects emotions, enabling natural communication.
[0551] hardware
[0552] The system primarily uses the following hardware:
[0553] 1. Device: An electronic device equipped with a camera or sensors (e.g., smartphone, tablet, etc.).
[0554] 2. Camera: A high-resolution camera that captures sign language movements in real time.
[0555] software
[0556] The system uses the following software and technologies:
[0557] 1. Image processing technology: Use OpenCV or similar tools to analyze sign language movements.
[0558] 2. Machine learning algorithm: TensorFlow or PyTorch is used to recognize specific words from behavioral data.
[0559] 3. Natural Language Processing (NLP) Technology: Generative AI models such as GPT-3 are used to generate sentence candidates based on recognized words.
[0560] 4. Emotion Recognition Technology: Use Amazon Rekognition or Azure Face API to identify the user's emotional state and adjust suggested text accordingly.
[0561] Data processing and calculation
[0562] Capture and transmit sign language movements
[0563] The user performs sign language gestures in front of the device. The device uses its built-in camera to capture the sign language gestures in real time and sends the video data frame by frame to the server. HTTP or WebSocket protocol is used for transmission.
[0564] Video data analysis and word recognition
[0565] The server analyzes the received video data using image processing technologies such as OpenCV and TensorFlow, as well as machine learning algorithms, to recognize specific words. It detects hand shapes and movement patterns and converts them into corresponding words based on a pre-trained model.
[0566] Sentence candidate generation
[0567] The server generates sentence candidates using a generative AI model such as GPT-3 based on recognized words. It takes a prompt message as input and generates and lists multiple sentence candidates. For example, if the user inputs the words "I," "go," and "school," the server will use a prompt message like the following:
[0568] Sign language words to input: I, go, school
[0569] Emotional state: Happy
[0570] Generation example:
[0571] I look forward to going to school.
[0572] I can't wait to go to school.
[0573] Identifying and adjusting emotional states
[0574] The server uses emotion recognition technologies such as Amazon Rekognition and Azure Face API to analyze the user's facial expressions and voice tone from video data and identify their emotional state. For example, if a user is smiling while using sign language, the server recognizes this as "happy." Based on the identified emotional state, the server adjusts the tone and content of the generated text suggestions.
[0575] Presentation and selection of sentence options
[0576] The generated sentence suggestions are sent to the device, which displays them on its screen. The user selects the desired sentence using touch gestures or finger movements. For example, "I'm looking forward to going to school" and "I can't wait to go to school" might be displayed, and the user selects from among them.
[0577] Display of final output
[0578] The device ultimately displays the text selected by the user. This allows the user to accurately convey what they want to communicate, including their emotions, to others.
[0579] As described above, the system of the present invention enables natural communication that reflects the user's sign language movements and emotions.
[0580] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0581] Step 1: The user performs a sign language action.
[0582] The user performs sign language actions in front of the device's camera. For example, they might perform sign language actions for "I," "go," and "school." The user's sign language actions are the input and are captured as video data.
[0583] Step 2: The device captures sign language movements and sends the video data to the server.
[0584] The device uses a built-in high-resolution camera to capture the user's sign language movements in real time. The captured video data is acquired frame by frame and transmitted to the server via Wi-Fi or mobile data communication. The input is video data of the sign language movements, and the output is video data transferred to the server.
[0585] Step 3: The server analyzes the video data and converts sign language actions into words.
[0586] The server uses image processing techniques such as OpenCV and TensorFlow / PyTorch, along with machine learning algorithms, to analyze the received video data. It detects hand shapes and movement patterns and uses a pre-trained model to convert those movements into corresponding specific words. The input is video data, and the output is a list of recognized words.
[0587] Step 4: The server generates sentence suggestions based on the words.
[0588] The server generates sentence candidates using a generative AI model (e.g., GPT-3) based on recognized word combinations. A prompt sentence is entered, and multiple sentence candidates are generated and listed. The input consists of a list of recognized words and a prompt sentence, and the output is a list of sentence candidates. For example, if the words "I," "go," and "school" are entered, the server generates sentence candidates such as "I look forward to going to school" and "I can't wait to go to school."
[0589] Step 5: The server uses the emotion engine to identify the emotional state and adjust the suggested sentences.
[0590] The server uses emotion recognition technology (such as Amazon Rekognition or Azure Face API) to analyze facial expressions and voice tone to identify the user's emotional state. For example, if the user is smiling and using sign language in video data, it will recognize that the user is "happy." The input is the user's video data, and the output is the identified emotional state. Subsequently, the tone and content of the generated sentence candidates are adjusted according to the emotional state. For example, if the user is happy, positive sentences will be prioritized.
[0591] Step 6: The server sends the adjusted text candidates to the terminal.
[0592] The server sends the adjusted sentence candidates to the terminal in JSON format. The input is the adjusted sentence candidates, and the output is the data sent to the terminal. Specifically, sentence candidates such as "I'm looking forward to going to school" and "I can't wait to go to school" are sent.
[0593] Step 7: The device presents text suggestions to the user, and the user makes a selection.
[0594] The device displays received text suggestions on its screen. The user selects the desired text using the touchscreen or finger movements. Text suggestions are sent to the device as input, and the text selected by the user is returned as output. For example, this might involve displaying multiple sentences on the screen and the user selecting "I am looking forward to going to school."
[0595] Step 8: The device displays the final text.
[0596] The device ultimately displays the text selected by the user. The user selects the text as input, and the final displayed text is the output. For example, the text "I am looking forward to going to school" might be displayed on the screen. This display allows the user to accurately convey what they want to communicate to others.
[0597] (Application Example 2)
[0598] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0599] This invention aims to improve the efficiency and accuracy of communication in the workplace by converting sign language used by hearing-impaired workers into text in real time, while also considering the worker's emotional state during the text generation. Furthermore, it aims to reduce stress and anxiety experienced by workers by providing appropriate feedback and support that reflects their emotional state, thereby improving work safety and productivity.
[0600] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0601] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words and the user's emotional state, means for presenting the generated sentence candidates to the user, and means for the user to select a generated sentence candidate. This makes sign language communication more efficient and reflects the user's emotional state, enabling richer and more accurate communication.
[0602] "Means for capturing user actions" refers to devices that use cameras or sensors to acquire the physical movements performed by a user as digital data.
[0603] "Means for analyzing captured motion data and recognizing specific words" refers to a device or software that uses machine learning algorithms or image processing techniques to identify specific words or meanings based on user motion data.
[0604] "Means for generating sentence candidates based on recognized words and the user's emotional state" refers to a device or software that generates multiple sentence candidates using words recognized from the user's sign language actions and the user's emotional state analyzed by an emotion engine.
[0605] "Means for presenting generated sentence options to the user" refers to a display that visually shows the user the options for the generated sentences, or a device that provides audio notifications.
[0606] "Means for the user to select generated sentence candidates" refers to an interface for the user to select a desired sentence from the presented sentence candidates, and includes touchscreens and other input devices.
[0607] "Sign language" refers to a means of communication using hand and finger movements, which are physical actions used to convey specific words or meanings.
[0608] "Emotional state" refers to the user's emotional and mood state, and is identified by analyzing facial expressions, tone of voice, and other factors.
[0609] "Natural language processing technology" refers to technologies that enable computers to understand and generate human language, and includes algorithms for text analysis and generation.
[0610] "Emotional analysis technology" is a technology that analyzes a user's facial expressions, voice tone, body movements, etc., to identify the user's current emotional state.
[0611] The embodiments for carrying out this invention will be described in detail below.
[0612] First, the system program is generated. This program captures the user's sign language movements, analyzes the movement data, and recognizes specific words. It also incorporates an emotion engine that recognizes the user's emotional state, and includes emotional expressions in the generated sentences.
[0613] The system consists of the following main hardware and software components:
[0614] 1. Means for capturing user actions
[0615] Cameras and sensors are used to capture the user's hand and finger movements in real time. Specifically, webcams and dedicated motion capture devices are used.
[0616] 2. A means of analyzing captured motion data and recognizing specific words.
[0617] Machine learning algorithms and image processing techniques are used to analyze captured sign language movements. Specifically, libraries such as TensorFlow and OpenCV are used.
[0618] 3. A means of generating sentence candidates based on recognized words and the user's emotional state.
[0619] This system generates multiple text candidates using natural language processing (NLP) and emotion analysis techniques. Specifically, it employs an emotion engine (e.g., EmotionEngine) and an NLP model (e.g., NLPModel).
[0620] 4. Means for presenting generated text candidates to the user
[0621] The system uses a display to present the user with generated text suggestions. This can be done using a computer monitor or projector.
[0622] 5. Means for the user to select generated text suggestions
[0623] The user selects the desired text using a touchscreen or other interface. Specifically, touch panels and pointing devices are used.
[0624] Furthermore, specific examples are given below.
[0625] As an example, consider a case where a hearing-impaired factory worker communicates in sign language that "the machine is broken." First, a camera captures the sign language movements, and the system analyzes the movements to recognize that "the machine is broken." Simultaneously, an emotion engine analyzes the worker's facial expressions to read their feelings of anxiety. Based on this, the system generates a suggested sentence, "The machine is broken. I am very anxious," and displays it on the screen. The worker can then select this sentence using a touch panel to inform other workers or supervisors.
[0626] Examples of prompt statements include the following:
[0627] "Imagine a system that uses sign language to report the situation within a factory. For example, if a worker expresses in sign language that 'the machine is broken' and shows an anxious expression, build a system that analyzes the sign language and emotions to provide appropriate feedback. This system combines sign language recognition and emotion recognition technologies to generate a detailed report."
[0628] Thus, this invention streamlines communication using sign language and reflects the emotions of the workers, thereby achieving more natural and enriching communication.
[0629] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0630] Step 1:
[0631] Camera-based motion capture
[0632] The server uses a camera to capture the user's sign language movements in real time. The input is the user's hand and finger movements, and the output is video frame data. This frame data is then sent to the next analysis step.
[0633] Step 2:
[0634] Sign language motion recognition
[0635] The server analyzes captured video frame data using image processing techniques and machine learning algorithms to recognize specific words. At this stage, TensorFlow and OpenCV are used to analyze hand shapes and movement patterns. The input is video frame data, and the output is the recognized word.
[0636] Step 3:
[0637] Emotion analysis
[0638] The server analyzes the user's facial expressions and voice tone from the captured video frame data and uses an EmotionEngine to identify the user's emotional state. The input is the same video frame data, and the output is the analyzed emotional state.
[0639] Step 4:
[0640] Sentence generation
[0641] The server generates multiple sentence candidates using natural language processing (NLP) techniques based on recognized words and analyzed sentiment states. An NLP model is used at this stage. The input is recognized words and sentiment states, and the output is multiple sentence candidates.
[0642] Step 5:
[0643] Suggestion of sentence candidates
[0644] The terminal visually displays the generated sentence candidates on its screen and presents them to the user. The input consists of multiple sentence candidates, and the output includes the sentence candidates displayed on the screen. The user is then ready to make a selection.
[0645] Step 6:
[0646] Selection of sentence candidates
[0647] The user selects their desired sentence from the presented options. The device acquires the user's selection data using a touchscreen or pointing device. The input is the user's selection operation, and the output is the final sentence selected by the user.
[0648] Step 7:
[0649] Output of selected text
[0650] The device displays the final text selected by the user on its screen, confirming the intent of the communication. The input is the text selected by the user, and the output is the text that is finally displayed on the device. This ensures that the user's intent is conveyed to the other person.
[0651] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0652] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0653] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0654] [Third Embodiment]
[0655] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0656] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0657] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0658] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0659] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0660] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0661] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0662] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0663] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0664] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0665] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0666] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0667] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate and communicate text using sign language.
[0668] Users perform sign language actions through a device equipped with a camera and sensors. The device uses the camera to capture these actions in real time. The camera continuously acquires video data and transmits it to a server.
[0669] The server analyzes the received video data and converts sign language movements into specific words. This uses image processing techniques and machine learning algorithms. Specifically, the server recognizes the shape and movement patterns of the hands and identifies the corresponding words based on a pre-trained model.
[0670] Next, the server generates sentence candidates based on the recognized words. Natural language processing (NLP) techniques are used here. The server generates multiple sentence candidates based on the recognized word combinations and lists them.
[0671] The generated sentence candidates are sent to the terminal, which then visually presents them to the user. Specifically, the candidate sentences are displayed on the screen, allowing the user to select one. The user uses finger movements or touch gestures to select the desired sentence from the presented candidates.
[0672] The device detects the user's selection and displays the final selected sentence. This allows the user to easily confirm the message they want to convey.
[0673] For example, if a user performs the sign language actions for "I," "go," and "school," the server recognizes these as specific words and generates several sentence options, such as "I go to school," "It is I who goes to school," and "The place I go to is school." The user then selects their desired sentence from these options, and the terminal displays the selected sentence as the final output.
[0674] In this way, the present invention is a system that can streamline communication using sign language and significantly reduce the burden on sign language users.
[0675] The following describes the processing flow.
[0676] Step 1:
[0677] The user performs sign language actions. A device equipped with a camera and sensors captures these sign language actions in real time. Specifically, the device continuously acquires video data, accurately capturing the shape and movement of the hand.
[0678] Step 2:
[0679] The terminal sends the acquired video data to the server. The server receives this video data and begins analysis. At this stage, data preprocessing such as noise reduction and color correction may be performed.
[0680] Step 3:
[0681] The server analyzes video data and converts sign language movements into specific words. The server uses image processing technology to analyze the contours and movements of the hands and maps the recognized movements to corresponding words through machine learning algorithms. This identifies each element of the sign language (e.g., subject, predicate, object, etc.).
[0682] Step 4:
[0683] The server uses natural language processing (NLP) techniques to generate sentence candidates based on the words it recognizes. For example, from the words "I," "go," and "school," multiple sentences such as "I go to school" and "It is I who go to school" are created.
[0684] Step 5:
[0685] The server sends generated sentence candidates to the terminal, which then visually presents them to the user. Specifically, multiple candidate sentences are displayed on the screen in a list format. Each sentence candidate is identified using a number or highlighting.
[0686] Step 6:
[0687] The user selects their desired text from a list of suggested texts. The user makes the selection using finger movements or touch gestures. The device captures the user's selection actions and sends that information to the server.
[0688] Step 7:
[0689] The server receives the user's selection information and finally confirms the selected text. The selected text is then sent from the server to the terminal.
[0690] Step 8:
[0691] The device visually displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others.
[0692] Through the processing steps described above, the present invention realizes efficient and intuitive text generation and communication using sign language.
[0693] (Example 1)
[0694] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0695] While modern communication largely relies on text-based technologies, their use can be limited for users who primarily rely on sign language, such as those with hearing impairments. Therefore, there is a need for systems that can efficiently and intuitively generate text using sign language and transmit it directly to others. However, existing technologies have been insufficient in accurately recognizing and transcribing sign language movements, hindering smooth communication for sign language users. In particular, challenges existed in the accurate capture and analysis of sign language movements, and the subsequent natural text generation process.
[0696] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0697] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, and means for generating sentence candidates based on the recognized words. This enables the real-time capture of the user's sign language actions, transmission of the action data to the server, and analysis of the sign language actions using image processing technology and machine learning algorithms. Furthermore, by utilizing natural language processing technology to generate multiple sentence candidates and presenting them visually to the user, the user can easily select their desired sentence and confirm the final sentence. This enables sign language users, including those with hearing impairments, to communicate more efficiently and intuitively.
[0698] "Means for capturing user actions" refer to devices or methods for recording a user's sign language movements in real time, and generally involve the use of cameras or sensors.
[0699] "Means for analyzing captured motion data and recognizing specific words" refers to technologies for analyzing acquired video data of sign language movements and identifying specific words based on that analysis, and includes image processing technologies and machine learning algorithms.
[0700] "A means of generating sentence candidates based on recognized words" refers to a system that automatically generates appropriate sentences by combining words recognized through analysis, and utilizes natural language processing technology.
[0701] "Means for presenting generated text candidates to the user" refers to devices or methods for displaying text candidates in a way that the user can easily review, and primarily uses displays or monitors.
[0702] "Means for the user to select generated sentence candidates" refers to an interface for the user to choose their preferred sentence from among several presented candidates, and includes touchscreens, pointing devices, and the like.
[0703] "Means for displaying selected text" refers to devices or methods for visually showing the text ultimately selected by the user, and displays or screens are used for this purpose.
[0704] "Methods of using cameras and sensors for capture" refers to technologies that utilize cameras and various sensors to record detailed sign language movements.
[0705] "Image processing techniques for identifying hand shape and movement patterns" are methods for characterizing hand shape and movement by analyzing data acquired from cameras and sensors, and utilize image processing libraries, etc.
[0706] "Methods for recognizing words using trained machine learning algorithms" refers to technologies that automatically identify corresponding words from sign language gestures using machine learning models that have been trained on a large amount of data in advance.
[0707] "Methods for generating sentence candidates using natural language processing technology" refers to methods that utilize machine learning models to generate multiple natural-sounding sentences from recognized word combinations.
[0708] "Means for a terminal to send video data acquired by its camera to a server" refers to a technology that allows a terminal to send video data captured in real time to a server via communication means such as the internet.
[0709] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate and communicate text using sign language. Specific embodiments are described below.
[0710] Hardware and software configuration
[0711] Hardware:
[0712] 1. Camera: Used to capture the user's sign language movements. Generally, a high-resolution camera is used.
[0713] 2. Sensors: Used to help accurately determine the position of hands and fingers. Depth sensors are examples of such sensors.
[0714] 3. Terminal: A device used to process captured video data and display the results to the user. Smartphones and tablets are typical examples.
[0715] 4. Display: Built into the device and used to visually present text suggestions to the user.
[0716] software:
[0717] 1. Image processing technology: Using libraries such as OpenCV, we extract the shape and movement characteristics of the hand from video data captured by the camera.
[0718] 2. Machine Learning Algorithms: Using TensorFlow, PyTorch, etc., a model is constructed and applied to convert sign language actions into specific words based on the extracted features.
[0719] 3. Natural Language Processing (NLP): Using generative AI models such as GPT and BERT, natural-sounding sentence candidates are generated based on recognized words.
[0720] Specific description of the system's operation
[0721] The user stands in front of the device's camera to perform sign language actions. The device's camera captures the user's sign language actions in real time. The video data acquired by the camera is transmitted to a server via the internet. The server analyzes the received video data and uses image processing technology to identify hand shapes and movement patterns. Next, a machine learning algorithm is applied to convert the identified sign language actions into corresponding words.
[0722] Next, natural language processing technology is used to generate multiple sentence candidates based on the recognized words. The generated sentence candidates are sent from the server to the terminal and displayed on the terminal's screen. The user selects the desired sentence from the presented candidates using touch gestures or other methods. Finally, the selected sentence is displayed on the terminal's screen, allowing the user to show the content they want to communicate to others.
[0723] Specific example
[0724] As an example, consider a case where a user performs the sign language actions for "I," "go," and "school." The device's camera captures these actions and sends the data to a server. The server analyzes the video data and identifies the words "I," "go," and "school." Next, using natural language processing technology, it generates several sentence options, such as "I go to school," "It is I who goes to school," and "The place I go to is school." The generated sentence options are sent to the device and displayed on the screen. The user selects their desired sentence, and the final sentence is displayed on the screen.
[0725] Example of a prompt
[0726] Example: "Please write a prompt for a system that analyzes sign language movements and generates multiple sentence options based on them."
[0727] As described above, the system of the present invention can significantly improve communication using sign language and reduce the communication burden on sign language users.
[0728] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0729] Step 1: Capture the user's sign language movements
[0730] Input: User's sign language gestures
[0731] Specific operation: The user stands in front of the device's camera to perform sign language actions. The device's camera captures the sign language actions in real time and continuously acquires video data.
[0732] Output: Acquired video data
[0733] Step 2: Sending video data
[0734] Input: Acquired video data
[0735] Specific operation: The device sends the captured video data to the server. The data is streamed to the server in real time over the internet.
[0736] Output: Video data sent to the server
[0737] Step 3: Analyzing video data
[0738] Input: Video data sent to the server
[0739] Specific operation: The server analyzes the received video data. First, it uses image processing techniques (such as OpenCV) to extract features of the hand's shape and movement. Afterwards, it identifies feature quantities (hand shape, position, movement pattern, etc.).
[0740] Output: Feature data
[0741] Step 4: Word Recognition
[0742] Input: Feature data
[0743] Specific operation: The server applies machine learning algorithms (TensorFlow or PyTorch) based on the extracted features to convert sign language actions into corresponding words. A pre-trained model identifies the appropriate word from the sign language action.
[0744] Output: Recognized words
[0745] Step 5: Generating text candidates
[0746] Input: Recognized word
[0747] Specific operation: The server generates multiple sentence candidates using natural language processing (NLP) techniques (GPT and BERT) based on recognized words. It then automatically generates appropriate sentence combinations using a generative AI model.
[0748] Output: Sentence suggestion list
[0749] Step 6: Send the suggested message
[0750] Input: Sentence suggestion list
[0751] Specific operation: The server sends a list of generated sentence suggestions to the terminal. The data is sent to the terminal via the network.
[0752] Output: List of suggested sentences sent to the terminal
[0753] Step 7: Display candidate sentences
[0754] Input: List of suggested sentences sent to the terminal
[0755] Specific operation: The device displays received text suggestions on its screen. It displays a list so that the user can visually see multiple text suggestions.
[0756] Output: Suggested sentences displayed on the screen
[0757] Step 8: User Selection
[0758] Input: Suggested sentences displayed on the screen
[0759] Specific operation: The user selects the desired sentence from the presented sentence options using finger movements or touch operations.
[0760] Output: Selected text
[0761] Step 9: Display the final sentence
[0762] Input: Selected text
[0763] Specific operation: The device displays the final sentence selected by the user on its screen. This allows the user to show others the message they ultimately wanted to convey.
[0764] Output: The final sentence displayed on the screen
[0765] (Application Example 1)
[0766] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0767] Traditional factory robot operation typically involves keyboards or touch panels, limiting intuitive control methods. This posed a problem, particularly for workers whose primary means of communication is sign language, as the operation was complex and difficult to understand. Furthermore, the lack of sufficient means to improve work efficiency was also a challenge.
[0768] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0769] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words, means for presenting the generated sentence candidates to the user, means for the user to select a generated sentence candidate, means for the user to give instructions for robot operation using sign language, and means for robot operation to execute the recognized sign language instructions. This makes it possible to operate the robot intuitively and efficiently using sign language.
[0770] "Means for capturing user movements" refers to devices that use cameras or sensors to capture the user's body movements and acquire those movements as digital data.
[0771] "A means of analyzing captured motion data and recognizing specific words" refers to the process of analyzing acquired digital data and identifying words that correspond to specific sign language or actions intended by the user.
[0772] "Methods for generating sentence candidates based on recognized words" refers to algorithms that combine words identified through analysis to generate candidates that form natural-sounding sentences.
[0773] "Means of presenting generated sentence candidates to the user" refers to an interface that visually displays multiple generated sentences as candidates on a display or screen, allowing the user to select one.
[0774] "Means for the user to select generated sentence candidates" refers to the means by which the user selects the desired sentence from among multiple displayed sentence candidates.
[0775] "A means by which a user can give instructions to operate a robot using sign language" refers to an interface for transmitting operating instructions to a robot or machine through sign language actions.
[0776] A "robot operating means for executing instructions based on recognized sign language" is a control device that allows a robot to perform specific operations based on instructions recognized as sign language actions.
[0777] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate text using sign language and to operate robots and give instructions in a factory setting.
[0778] hardware
[0779] The server and terminal include the following hardware:
[0780] Camera (e.g., Logitech C920): Captures the user's sign language movements in real time.
[0781] Robot control device: A device for controlling a robot based on recognized sign language movements.
[0782] software
[0783] The following software will be used:
[0784] Python: A programming language used for the overall system implementation.
[0785] OpenCV (CV2): A library for capturing and pre-processing camera footage.
[0786] Keras: A machine learning framework for running sign language gesture recognition models.
[0787] Natural language processing libraries (e.g., spaCy, GPT model): These are libraries for generating text candidates based on sign language actions.
[0788] System operation
[0789] The server receives video data of sign language movements transmitted from the user's camera. The video data is first preprocessed using OpenCV and then fed into a Keras model. The Keras model uses a pre-trained machine learning algorithm to convert the sign language movements into corresponding words.
[0790] The converted words are then transformed into multiple sentence candidates using a natural language processing library (e.g., spaCy or the GPT model). These sentence candidates are sent from the server to the terminal and displayed visually on the terminal's screen.
[0791] The user selects their desired sentence from several sentence options displayed on the screen. This selection is made through touch operations or finger movements. The final selected sentence is displayed on the device and then sent to the robot.
[0792] The robot performs specific actions based on the selected sentence. For example, if the user performs the sign language action for "pick up the part and move it," the robot will receive instructions such as "pick up the part and move it 5 meters."
[0793] Specific example
[0794] When a worker performs a sign language gesture for "pick up the part and move it," the server analyzes it and recognizes it as a corresponding word. The server then uses natural language processing technology to generate sentence options such as "pick up the part and move it 5 meters" or "take the part and move it to the designated location." These options are sent to the terminal, and the user selects the one sentence they prefer. Finally, the robot performs the action based on the selected instruction.
[0795] Example of a prompt
[0796] Example of sign language recognition results for use in analyzing sign language movements and generating text: Please generate phrases corresponding to the sign language words "part," "pick up," and "move." Possible instruction sentences would be something like, "Pick up the part and move 5 meters."
[0797] This invention enables intuitive and efficient operation and instruction using sign language in factories. Furthermore, it makes operation easier and improves work efficiency for workers who primarily use sign language for communication.
[0798] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0799] Step 1:
[0800] The server receives video data of sign language movements transmitted from the user's camera in real time. This video data is data captured by the camera of the sign language movements performed by the user. The data input is a video stream from the camera, and the output is the frame data of that video. Specifically, the camera continuously captures the user's hand movements and sends each frame to the server as digital data.
[0801] Step 2:
[0802] The server analyzes the received video data. OpenCV is used for preprocessing the video. In this step, the necessary sign language movements are extracted from the video, and noise is removed. The input for this process is frame data, and the output is preprocessed image data of the sign language movements. Specifically, this involves image smoothing and edge detection to enhance the shape of the hands.
[0803] Step 3:
[0804] The server supplies pre-processed video data to a Keras model, which runs a machine learning algorithm that converts sign language gestures into corresponding words. In this step, the input is pre-processed image data, and the output is the recognized words. Specifically, the model analyzes certain features in the image and maps them to pre-trained words.
[0805] Step 4:
[0806] The server generates multiple sentence candidates using a natural language processing library based on the recognized words. In this step, the input is the recognized words, and the output is a list of sentence candidates. Specifically, an NLP library (e.g., spaCy or GPT model) analyzes the word combinations and generates candidates as meaningful sentences.
[0807] Step 5:
[0808] The terminal presents the user with sentence suggestions sent from the server. In this step, the input is a list of generated sentence suggestions, and the output is the sentence displayed on the user's screen. Specifically, the terminal displays the sentence suggestions on its screen so that the user can visually confirm them.
[0809] Step 6:
[0810] The user selects a desired sentence from several sentence options displayed on the screen. In this step, the input is the sentence options displayed on the screen, and the output is the sentence selected by the user. Specifically, the user uses a touchscreen or mouse to select the desired sentence.
[0811] Step 7:
[0812] The terminal sends the text selected by the user to the server, which then distributes it to the robot control unit. In this step, the input is the text selected by the user, and the output is the control instruction for the robot. Specifically, the terminal sends the selected text to the server, and the server relays that instruction to the robot.
[0813] Step 8:
[0814] The robot performs specific operations based on the instructions it receives. In this step, the input is the instruction sent to the robot, and the output is the robot's physical movement. For example, based on the instruction "pick up and move the part," the robot arm picks up the part and moves it to the designated location.
[0815] This system enables intuitive robot operation using sign language, leading to increased efficiency in factory operations.
[0816] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0817] This invention combines a system that captures a user's sign language movements, analyzes the movement data to recognize specific words, with an emotion engine that recognizes the user's emotions. This system enables users to efficiently and intuitively generate sentences using sign language, and in addition, to generate sentences that include appropriate emotional expressions.
[0818] The user performs sign language actions through a device equipped with a camera and sensors. The device uses the camera to capture these sign language actions in real time. The camera continuously acquires video data and transmits it to a server.
[0819] The server analyzes the received video data and converts sign language movements into specific words. This uses image processing techniques and machine learning algorithms. Specifically, the server recognizes the shape and movement patterns of the hands and identifies the corresponding words based on a pre-trained model.
[0820] Next, the server generates sentence candidates based on the recognized words. Natural language processing (NLP) techniques are used here. The server generates multiple sentence candidates based on the recognized word combinations and lists them.
[0821] Furthermore, this system incorporates an emotion engine. The emotion engine analyzes the user's facial expressions and tone of voice to identify the user's emotional state. For example, if a user is smiling while performing sign language, the emotion engine recognizes that the user is "happy."
[0822] Based on the emotional state identified by the emotion engine, the server adjusts the tone and content of the generated text suggestions. For example, if the user is "happy," the server prioritizes generating text suggestions with a positive tone. Conversely, if the user is "sad," the server generates text suggestions that are more empathetic to that emotion.
[0823] The generated sentence candidates are sent to the terminal, which then visually presents them to the user. Specifically, the candidate sentences are displayed on the screen, allowing the user to select one. The user uses finger movements or touch gestures to select the desired sentence from the presented candidates.
[0824] The device detects the user's selection and displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others. The introduction of an emotion engine is expected to result in a more accurate and richer expression of the message.
[0825] For example, if a user performs the sign language actions for "I," "go," and "school," and the emotion engine recognizes the user's emotional state as "happy," the system will generate positive sentence options such as "I'm looking forward to going to school" or "I can't wait to go to school." The user then selects their desired sentence from these options, and the device displays the selected sentence as the final output.
[0826] In this way, the present invention is a system that can streamline communication using sign language and reflect the emotions of sign language users, thereby realizing more natural and enriching communication.
[0827] The following describes the processing flow.
[0828] Step 1:
[0829] The user performs sign language actions. A device equipped with a camera and sensors captures these sign language actions in real time. Specifically, the device uses its camera to continuously acquire video data, accurately capturing the shape and movement of the hand.
[0830] Step 2:
[0831] The terminal sends the acquired video data to the server. The server receives this video data and begins analysis. Data compression technology is used during data transfer for efficient transmission.
[0832] Step 3:
[0833] The server analyzes video data and converts sign language movements into specific words. The server uses image processing technology to analyze the contours and movements of the hands and maps the recognized movements to corresponding words through machine learning algorithms. Specifically, the server analyzes the shape, position, and movement patterns of the hands to recognize words such as "subject," "predicate," and "object."
[0834] Step 4:
[0835] The server uses natural language processing (NLP) techniques to generate sentence candidates based on the recognized words. For example, from the recognized words "I," "go," and "school," several sentences such as "I go to school" and "It is I who go to school" are created. The server then lists these sentence candidates.
[0836] Step 5:
[0837] The emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotions. The device uses its camera and microphone to capture the user's facial expressions and voice data, and sends this data to the server.
[0838] Step 6:
[0839] The server analyzes the facial and voice data it receives to identify the user's emotional state. The emotion engine uses facial recognition and voice analysis technologies to recognize the user's emotional state, such as "happy," "sad," or "surprised."
[0840] Step 7:
[0841] The server adjusts the tone and content of the generated sentence suggestions based on the identified emotional state. For example, if the user is perceived as "happy," the server prioritizes generating sentence suggestions with a positive tone. Specifically, it might generate sentence suggestions that include emotional expressions such as, "I'm looking forward to going to school."
[0842] Step 8:
[0843] The server sends generated sentiment-reflecting sentence candidates to the terminal, which then visually presents them to the user. Multiple candidate sentences are displayed on the screen in a list format, with each candidate numbered or highlighted.
[0844] Step 9:
[0845] The user selects their desired text from a list of suggested texts. The user makes the selection using finger movements or touch gestures. The device captures the user's selection actions and sends that information to the server.
[0846] Step 10:
[0847] The server receives the user's selection information and finally confirms the selected text. The selected text is then sent from the server to the terminal.
[0848] Step 11:
[0849] The device visually displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others.
[0850] Through the processing steps described above, the present invention realizes efficient and intuitive text generation and communication using sign language. Furthermore, the introduction of an emotion engine enables natural and rich communication that also reflects emotional expression.
[0851] (Example 2)
[0852] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0853] Conventional sign language recognition systems only recognize sign language movements, and the generated text does not reflect the user's emotions, resulting in the problem of inaccurately conveying the intended message. Furthermore, providing a diverse range of example sentences during text generation is difficult, limiting communication possibilities.
[0854] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words, means for presenting the generated sentence candidates to the user, means for the user to select a generated sentence candidate, and means for recognizing the user's emotional state and adjusting the tone and content of the sentence candidates based on that emotional state. This makes it possible to generate accurate sentences based on the user's sign language actions and to provide a variety of sentence candidates that reflect the user's emotions.
[0855] A "user" is a person who uses this system and is the entity that inputs data through sign language actions.
[0856] "Means for capturing movements" refers to devices or technologies that use cameras or sensors to acquire a user's sign language movements in real time.
[0857] "Motion data" refers to video information and sensor data that digitally records the user's sign language movements.
[0858] "Means of analysis" refers to methods and techniques for analyzing acquired motion data using machine learning algorithms and image processing technologies, and converting it into specific words.
[0859] A "specific word" is a string of characters with meaning recognized from the analyzed motion data, representing a conversion of sign language movements into language.
[0860] "Methods for generating sentence candidates" refer to methods and techniques for creating multiple sentences based on recognized words using natural language processing technology or generative AI models.
[0861] "Means for presenting generated sentence candidates to the user" refers to a technology that visually displays multiple sentence candidates to the user using a display device such as a screen.
[0862] "Means for users to select generated text candidates" refers to a device or technology that allows a user to select a desired text from a visually displayed list of text candidates using a touchscreen or finger movements.
[0863] "Means of recognizing emotional states" refers to methods and technologies that use facial recognition technology or voice analysis technology to identify a user's emotions.
[0864] "Means of adjusting tone and content" refers to methods and techniques for appropriately modifying the style and content of candidate texts generated based on the recognized emotional state of the user.
[0865] This invention relates to a system that captures a user's sign language movements, analyzes the movement data to recognize specific words, and further recognizes the user's emotions to generate sentence candidates based on those emotions. This system efficiently and intuitively generates sentences based on the user's sign language movements and reflects emotions, enabling natural communication.
[0866] hardware
[0867] The system primarily uses the following hardware:
[0868] 1. Device: An electronic device equipped with a camera or sensors (e.g., smartphone, tablet, etc.).
[0869] 2. Camera: A high-resolution camera that captures sign language movements in real time.
[0870] software
[0871] The system uses the following software and technologies:
[0872] 1. Image processing technology: Use OpenCV or similar tools to analyze sign language movements.
[0873] 2. Machine learning algorithm: TensorFlow or PyTorch is used to recognize specific words from behavioral data.
[0874] 3. Natural Language Processing (NLP) Technology: Generative AI models such as GPT-3 are used to generate sentence candidates based on recognized words.
[0875] 4. Emotion Recognition Technology: Use Amazon Rekognition or Azure Face API to identify the user's emotional state and adjust suggested text accordingly.
[0876] Data processing and calculation
[0877] Capture and transmit sign language movements
[0878] The user performs sign language gestures in front of the device. The device uses its built-in camera to capture the sign language gestures in real time and sends the video data frame by frame to the server. HTTP or WebSocket protocol is used for transmission.
[0879] Video data analysis and word recognition
[0880] The server analyzes the received video data using image processing technologies such as OpenCV and TensorFlow, as well as machine learning algorithms, to recognize specific words. It detects hand shapes and movement patterns and converts them into corresponding words based on a pre-trained model.
[0881] Sentence candidate generation
[0882] The server generates sentence candidates using a generative AI model such as GPT-3 based on recognized words. It takes a prompt message as input and generates and lists multiple sentence candidates. For example, if the user inputs the words "I," "go," and "school," the server will use a prompt message like the following:
[0883] Sign language words to input: I, go, school
[0884] Emotional state: Happy
[0885] Generation example:
[0886] I look forward to going to school.
[0887] I can't wait to go to school.
[0888] Identifying and adjusting emotional states
[0889] The server uses emotion recognition technologies such as Amazon Rekognition and Azure Face API to analyze the user's facial expressions and voice tone from video data and identify their emotional state. For example, if a user is smiling while using sign language, the server recognizes this as "happy." Based on the identified emotional state, the server adjusts the tone and content of the generated text suggestions.
[0890] Presentation and selection of sentence options
[0891] The generated sentence suggestions are sent to the device, which displays them on its screen. The user selects the desired sentence using touch gestures or finger movements. For example, "I'm looking forward to going to school" and "I can't wait to go to school" might be displayed, and the user selects from among them.
[0892] Display of final output
[0893] The device ultimately displays the text selected by the user. This allows the user to accurately convey what they want to communicate, including their emotions, to others.
[0894] As described above, the system of the present invention enables natural communication that reflects the user's sign language movements and emotions.
[0895] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0896] Step 1: The user performs a sign language action.
[0897] The user performs sign language actions in front of the device's camera. For example, they might perform sign language actions for "I," "go," and "school." The user's sign language actions are the input and are captured as video data.
[0898] Step 2: The device captures sign language movements and sends the video data to the server.
[0899] The device uses a built-in high-resolution camera to capture the user's sign language movements in real time. The captured video data is acquired frame by frame and transmitted to the server via Wi-Fi or mobile data communication. The input is video data of the sign language movements, and the output is video data transferred to the server.
[0900] Step 3: The server analyzes the video data and converts sign language actions into words.
[0901] The server uses image processing techniques such as OpenCV and TensorFlow / PyTorch, along with machine learning algorithms, to analyze the received video data. It detects hand shapes and movement patterns and uses a pre-trained model to convert those movements into corresponding specific words. The input is video data, and the output is a list of recognized words.
[0902] Step 4: The server generates sentence suggestions based on the words.
[0903] The server generates sentence candidates using a generative AI model (e.g., GPT-3) based on recognized word combinations. A prompt sentence is entered, and multiple sentence candidates are generated and listed. The input consists of a list of recognized words and a prompt sentence, and the output is a list of sentence candidates. For example, if the words "I," "go," and "school" are entered, the server generates sentence candidates such as "I look forward to going to school" and "I can't wait to go to school."
[0904] Step 5: The server uses the emotion engine to identify the emotional state and adjust the suggested sentences.
[0905] The server uses emotion recognition technology (such as Amazon Rekognition or Azure Face API) to analyze facial expressions and voice tone to identify the user's emotional state. For example, if the user is smiling and using sign language in video data, it will recognize that the user is "happy." The input is the user's video data, and the output is the identified emotional state. Subsequently, the tone and content of the generated sentence candidates are adjusted according to the emotional state. For example, if the user is happy, positive sentences will be prioritized.
[0906] Step 6: The server sends the adjusted text candidates to the terminal.
[0907] The server sends the adjusted sentence candidates to the terminal in JSON format. The input is the adjusted sentence candidates, and the output is the data sent to the terminal. Specifically, sentence candidates such as "I'm looking forward to going to school" and "I can't wait to go to school" are sent.
[0908] Step 7: The device presents text suggestions to the user, and the user makes a selection.
[0909] The device displays received text suggestions on its screen. The user selects the desired text using the touchscreen or finger movements. Text suggestions are sent to the device as input, and the text selected by the user is returned as output. For example, this might involve displaying multiple sentences on the screen and the user selecting "I am looking forward to going to school."
[0910] Step 8: The device displays the final text.
[0911] The device ultimately displays the text selected by the user. The user selects the text as input, and the final displayed text is the output. For example, the text "I am looking forward to going to school" might be displayed on the screen. This display allows the user to accurately convey what they want to communicate to others.
[0912] (Application Example 2)
[0913] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0914] This invention aims to improve the efficiency and accuracy of communication in the workplace by converting sign language used by hearing-impaired workers into text in real time, while also considering the worker's emotional state during the text generation. Furthermore, it aims to reduce stress and anxiety experienced by workers by providing appropriate feedback and support that reflects their emotional state, thereby improving work safety and productivity.
[0915] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0916] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words and the user's emotional state, means for presenting the generated sentence candidates to the user, and means for the user to select a generated sentence candidate. This makes sign language communication more efficient and reflects the user's emotional state, enabling richer and more accurate communication.
[0917] "Means for capturing user actions" refers to devices that use cameras or sensors to acquire the physical movements performed by a user as digital data.
[0918] "Means for analyzing captured motion data and recognizing specific words" refers to a device or software that uses machine learning algorithms or image processing techniques to identify specific words or meanings based on user motion data.
[0919] "Means for generating sentence candidates based on recognized words and the user's emotional state" refers to a device or software that generates multiple sentence candidates using words recognized from the user's sign language actions and the user's emotional state analyzed by an emotion engine.
[0920] "Means for presenting generated sentence options to the user" refers to a display that visually shows the user the options for the generated sentences, or a device that provides audio notifications.
[0921] "Means for the user to select generated sentence candidates" refers to an interface for the user to select a desired sentence from the presented sentence candidates, and includes touchscreens and other input devices.
[0922] "Sign language" refers to a means of communication using hand and finger movements, which are physical actions used to convey specific words or meanings.
[0923] "Emotional state" refers to the user's emotional and mood state, and is identified by analyzing facial expressions, tone of voice, and other factors.
[0924] "Natural language processing technology" refers to technologies that enable computers to understand and generate human language, and includes algorithms for text analysis and generation.
[0925] "Emotional analysis technology" is a technology that analyzes a user's facial expressions, voice tone, body movements, etc., to identify the user's current emotional state.
[0926] The embodiments for carrying out this invention will be described in detail below.
[0927] First, the system program is generated. This program captures the user's sign language movements, analyzes the movement data, and recognizes specific words. It also incorporates an emotion engine that recognizes the user's emotional state, and includes emotional expressions in the generated sentences.
[0928] The system consists of the following main hardware and software components:
[0929] 1. Means for capturing user actions
[0930] Cameras and sensors are used to capture the user's hand and finger movements in real time. Specifically, webcams and dedicated motion capture devices are used.
[0931] 2. A means of analyzing captured motion data and recognizing specific words.
[0932] Machine learning algorithms and image processing techniques are used to analyze captured sign language movements. Specifically, libraries such as TensorFlow and OpenCV are used.
[0933] 3. A means of generating sentence candidates based on recognized words and the user's emotional state.
[0934] This system generates multiple text candidates using natural language processing (NLP) and emotion analysis techniques. Specifically, it employs an emotion engine (e.g., EmotionEngine) and an NLP model (e.g., NLPModel).
[0935] 4. Means for presenting generated text candidates to the user
[0936] The system uses a display to present the user with generated text suggestions. This can be done using a computer monitor or projector.
[0937] 5. Means for the user to select generated text suggestions
[0938] The user selects the desired text using a touchscreen or other interface. Specifically, touch panels and pointing devices are used.
[0939] Furthermore, specific examples are given below.
[0940] As an example, consider a case where a hearing-impaired factory worker communicates in sign language that "the machine is broken." First, a camera captures the sign language movements, and the system analyzes the movements to recognize that "the machine is broken." Simultaneously, an emotion engine analyzes the worker's facial expressions to read their feelings of anxiety. Based on this, the system generates a suggested sentence, "The machine is broken. I am very anxious," and displays it on the screen. The worker can then select this sentence using a touch panel to inform other workers or supervisors.
[0941] Examples of prompt statements include the following:
[0942] "Imagine a system that uses sign language to report the situation within a factory. For example, if a worker expresses in sign language that 'the machine is broken' and shows an anxious expression, build a system that analyzes the sign language and emotions to provide appropriate feedback. This system combines sign language recognition and emotion recognition technologies to generate a detailed report."
[0943] Thus, this invention streamlines communication using sign language and reflects the emotions of the workers, thereby achieving more natural and enriching communication.
[0944] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0945] Step 1:
[0946] Camera-based motion capture
[0947] The server uses a camera to capture the user's sign language movements in real time. The input is the user's hand and finger movements, and the output is video frame data. This frame data is then sent to the next analysis step.
[0948] Step 2:
[0949] Sign language motion recognition
[0950] The server analyzes captured video frame data using image processing techniques and machine learning algorithms to recognize specific words. At this stage, TensorFlow and OpenCV are used to analyze hand shapes and movement patterns. The input is video frame data, and the output is the recognized word.
[0951] Step 3:
[0952] Emotion analysis
[0953] The server analyzes the user's facial expressions and voice tone from the captured video frame data and uses an EmotionEngine to identify the user's emotional state. The input is the same video frame data, and the output is the analyzed emotional state.
[0954] Step 4:
[0955] Sentence generation
[0956] The server generates multiple sentence candidates using natural language processing (NLP) techniques based on recognized words and analyzed sentiment states. An NLP model is used at this stage. The input is recognized words and sentiment states, and the output is multiple sentence candidates.
[0957] Step 5:
[0958] Suggestion of sentence candidates
[0959] The terminal visually displays the generated sentence candidates on its screen and presents them to the user. The input consists of multiple sentence candidates, and the output includes the sentence candidates displayed on the screen. The user is then ready to make a selection.
[0960] Step 6:
[0961] Selection of sentence candidates
[0962] The user selects their desired sentence from the presented options. The device acquires the user's selection data using a touchscreen or pointing device. The input is the user's selection operation, and the output is the final sentence selected by the user.
[0963] Step 7:
[0964] Output of selected text
[0965] The device displays the final text selected by the user on its screen, confirming the intent of the communication. The input is the text selected by the user, and the output is the text that is finally displayed on the device. This ensures that the user's intent is conveyed to the other person.
[0966] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0967] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0968] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0969] [Fourth Embodiment]
[0970] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0971] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0972] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0973] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0974] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0975] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0976] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0977] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0978] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0979] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0980] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0981] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0982] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0983] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate and communicate text using sign language.
[0984] Users perform sign language actions through a device equipped with a camera and sensors. The device uses the camera to capture these actions in real time. The camera continuously acquires video data and transmits it to a server.
[0985] The server analyzes the received video data and converts sign language movements into specific words. This uses image processing techniques and machine learning algorithms. Specifically, the server recognizes the shape and movement patterns of the hands and identifies the corresponding words based on a pre-trained model.
[0986] Next, the server generates sentence candidates based on the recognized words. Natural language processing (NLP) techniques are used here. The server generates multiple sentence candidates based on the recognized word combinations and lists them.
[0987] The generated sentence candidates are sent to the terminal, which then visually presents them to the user. Specifically, the candidate sentences are displayed on the screen, allowing the user to select one. The user uses finger movements or touch gestures to select the desired sentence from the presented candidates.
[0988] The device detects the user's selection and displays the final selected sentence. This allows the user to easily confirm the message they want to convey.
[0989] For example, if a user performs the sign language actions for "I," "go," and "school," the server recognizes these as specific words and generates several sentence options, such as "I go to school," "It is I who goes to school," and "The place I go to is school." The user then selects their desired sentence from these options, and the terminal displays the selected sentence as the final output.
[0990] In this way, the present invention is a system that can streamline communication using sign language and significantly reduce the burden on sign language users.
[0991] The following describes the processing flow.
[0992] Step 1:
[0993] The user performs sign language actions. A device equipped with a camera and sensors captures these sign language actions in real time. Specifically, the device continuously acquires video data, accurately capturing the shape and movement of the hand.
[0994] Step 2:
[0995] The terminal sends the acquired video data to the server. The server receives this video data and begins analysis. At this stage, data preprocessing such as noise reduction and color correction may be performed.
[0996] Step 3:
[0997] The server analyzes video data and converts sign language movements into specific words. The server uses image processing technology to analyze the contours and movements of the hands and maps the recognized movements to corresponding words through machine learning algorithms. This identifies each element of the sign language (e.g., subject, predicate, object, etc.).
[0998] Step 4:
[0999] The server uses natural language processing (NLP) techniques to generate sentence candidates based on the words it recognizes. For example, from the words "I," "go," and "school," multiple sentences such as "I go to school" and "It is I who go to school" are created.
[1000] Step 5:
[1001] The server sends generated sentence candidates to the terminal, which then visually presents them to the user. Specifically, multiple candidate sentences are displayed on the screen in a list format. Each sentence candidate is identified using a number or highlighting.
[1002] Step 6:
[1003] The user selects their desired text from a list of suggested texts. The user makes the selection using finger movements or touch gestures. The device captures the user's selection actions and sends that information to the server.
[1004] Step 7:
[1005] The server receives the user's selection information and finally confirms the selected text. The selected text is then sent from the server to the terminal.
[1006] Step 8:
[1007] The device visually displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others.
[1008] Through the processing steps described above, the present invention realizes efficient and intuitive text generation and communication using sign language.
[1009] (Example 1)
[1010] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1011] While modern communication largely relies on text-based technologies, their use can be limited for users who primarily rely on sign language, such as those with hearing impairments. Therefore, there is a need for systems that can efficiently and intuitively generate text using sign language and transmit it directly to others. However, existing technologies have been insufficient in accurately recognizing and transcribing sign language movements, hindering smooth communication for sign language users. In particular, challenges existed in the accurate capture and analysis of sign language movements, and the subsequent natural text generation process.
[1012] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1013] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, and means for generating sentence candidates based on the recognized words. This enables the real-time capture of the user's sign language actions, transmission of the action data to the server, and analysis of the sign language actions using image processing technology and machine learning algorithms. Furthermore, by utilizing natural language processing technology to generate multiple sentence candidates and presenting them visually to the user, the user can easily select their desired sentence and confirm the final sentence. This enables sign language users, including those with hearing impairments, to communicate more efficiently and intuitively.
[1014] "Means for capturing user actions" refer to devices or methods for recording a user's sign language movements in real time, and generally involve the use of cameras or sensors.
[1015] "Means for analyzing captured motion data and recognizing specific words" refers to technologies for analyzing acquired video data of sign language movements and identifying specific words based on that analysis, and includes image processing technologies and machine learning algorithms.
[1016] "A means of generating sentence candidates based on recognized words" refers to a system that automatically generates appropriate sentences by combining words recognized through analysis, and utilizes natural language processing technology.
[1017] "Means for presenting generated text candidates to the user" refers to devices or methods for displaying text candidates in a way that the user can easily review, and primarily uses displays or monitors.
[1018] "Means for the user to select generated sentence candidates" refers to an interface for the user to choose their preferred sentence from among several presented candidates, and includes touchscreens, pointing devices, and the like.
[1019] "Means for displaying selected text" refers to devices or methods for visually showing the text ultimately selected by the user, and displays or screens are used for this purpose.
[1020] "Methods of using cameras and sensors for capture" refers to technologies that utilize cameras and various sensors to record detailed sign language movements.
[1021] "Image processing techniques for identifying hand shape and movement patterns" are methods for characterizing hand shape and movement by analyzing data acquired from cameras and sensors, and utilize image processing libraries, etc.
[1022] "Methods for recognizing words using trained machine learning algorithms" refers to technologies that automatically identify corresponding words from sign language gestures using machine learning models that have been trained on a large amount of data in advance.
[1023] "Methods for generating sentence candidates using natural language processing technology" refers to methods that utilize machine learning models to generate multiple natural-sounding sentences from recognized word combinations.
[1024] "Means for a terminal to send video data acquired by its camera to a server" refers to a technology that allows a terminal to send video data captured in real time to a server via communication means such as the internet.
[1025] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate and communicate text using sign language. Specific embodiments are described below.
[1026] Hardware and software configuration
[1027] Hardware:
[1028] 1. Camera: Used to capture the user's sign language movements. Generally, a high-resolution camera is used.
[1029] 2. Sensors: Used to help accurately determine the position of hands and fingers. Depth sensors are examples of such sensors.
[1030] 3. Terminal: A device used to process captured video data and display the results to the user. Smartphones and tablets are typical examples.
[1031] 4. Display: Built into the device and used to visually present text suggestions to the user.
[1032] software:
[1033] 1. Image processing technology: Using libraries such as OpenCV, we extract the shape and movement characteristics of the hand from video data captured by the camera.
[1034] 2. Machine Learning Algorithms: Using TensorFlow, PyTorch, etc., a model is constructed and applied to convert sign language actions into specific words based on the extracted features.
[1035] 3. Natural Language Processing (NLP): Using generative AI models such as GPT and BERT, natural-sounding sentence candidates are generated based on recognized words.
[1036] Specific description of the system's operation
[1037] The user stands in front of the device's camera to perform sign language actions. The device's camera captures the user's sign language actions in real time. The video data acquired by the camera is transmitted to a server via the internet. The server analyzes the received video data and uses image processing technology to identify hand shapes and movement patterns. Next, a machine learning algorithm is applied to convert the identified sign language actions into corresponding words.
[1038] Next, natural language processing technology is used to generate multiple sentence candidates based on the recognized words. The generated sentence candidates are sent from the server to the terminal and displayed on the terminal's screen. The user selects the desired sentence from the presented candidates using touch gestures or other methods. Finally, the selected sentence is displayed on the terminal's screen, allowing the user to show the content they want to communicate to others.
[1039] Specific example
[1040] As an example, consider a case where a user performs the sign language actions for "I," "go," and "school." The device's camera captures these actions and sends the data to a server. The server analyzes the video data and identifies the words "I," "go," and "school." Next, using natural language processing technology, it generates several sentence options, such as "I go to school," "It is I who goes to school," and "The place I go to is school." The generated sentence options are sent to the device and displayed on the screen. The user selects their desired sentence, and the final sentence is displayed on the screen.
[1041] Example of a prompt
[1042] Example: "Please write a prompt for a system that analyzes sign language movements and generates multiple sentence options based on them."
[1043] As described above, the system of the present invention can significantly improve communication using sign language and reduce the communication burden on sign language users.
[1044] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1045] Step 1: Capture the user's sign language movements
[1046] Input: User's sign language gestures
[1047] Specific operation: The user stands in front of the device's camera to perform sign language actions. The device's camera captures the sign language actions in real time and continuously acquires video data.
[1048] Output: Acquired video data
[1049] Step 2: Sending video data
[1050] Input: Acquired video data
[1051] Specific operation: The device sends the captured video data to the server. The data is streamed to the server in real time over the internet.
[1052] Output: Video data sent to the server
[1053] Step 3: Analyzing video data
[1054] Input: Video data sent to the server
[1055] Specific operation: The server analyzes the received video data. First, it uses image processing techniques (such as OpenCV) to extract features of the hand's shape and movement. Afterwards, it identifies feature quantities (hand shape, position, movement pattern, etc.).
[1056] Output: Feature data
[1057] Step 4: Word Recognition
[1058] Input: Feature data
[1059] Specific operation: The server applies machine learning algorithms (TensorFlow or PyTorch) based on the extracted features to convert sign language actions into corresponding words. A pre-trained model identifies the appropriate word from the sign language action.
[1060] Output: Recognized words
[1061] Step 5: Generating text candidates
[1062] Input: Recognized word
[1063] Specific operation: The server generates multiple sentence candidates using natural language processing (NLP) techniques (GPT and BERT) based on recognized words. It then automatically generates appropriate sentence combinations using a generative AI model.
[1064] Output: Sentence suggestion list
[1065] Step 6: Send the suggested message
[1066] Input: Sentence suggestion list
[1067] Specific operation: The server sends a list of generated sentence suggestions to the terminal. The data is sent to the terminal via the network.
[1068] Output: List of suggested sentences sent to the terminal
[1069] Step 7: Display candidate sentences
[1070] Input: List of suggested sentences sent to the terminal
[1071] Specific operation: The device displays received text suggestions on its screen. It displays a list so that the user can visually see multiple text suggestions.
[1072] Output: Suggested sentences displayed on the screen
[1073] Step 8: User Selection
[1074] Input: Suggested sentences displayed on the screen
[1075] Specific operation: The user selects the desired sentence from the presented sentence options using finger movements or touch operations.
[1076] Output: Selected text
[1077] Step 9: Display the final sentence
[1078] Input: Selected text
[1079] Specific operation: The device displays the final sentence selected by the user on its screen. This allows the user to show others the message they ultimately wanted to convey.
[1080] Output: The final sentence displayed on the screen
[1081] (Application Example 1)
[1082] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1083] Traditional factory robot operation typically involves keyboards or touch panels, limiting intuitive control methods. This posed a problem, particularly for workers whose primary means of communication is sign language, as the operation was complex and difficult to understand. Furthermore, the lack of sufficient means to improve work efficiency was also a challenge.
[1084] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1085] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words, means for presenting the generated sentence candidates to the user, means for the user to select a generated sentence candidate, means for the user to give instructions for robot operation using sign language, and means for robot operation to execute the recognized sign language instructions. This makes it possible to operate the robot intuitively and efficiently using sign language.
[1086] "Means for capturing user movements" refers to devices that use cameras or sensors to capture the user's body movements and acquire those movements as digital data.
[1087] "A means of analyzing captured motion data and recognizing specific words" refers to the process of analyzing acquired digital data and identifying words that correspond to specific sign language or actions intended by the user.
[1088] "Methods for generating sentence candidates based on recognized words" refers to algorithms that combine words identified through analysis to generate candidates that form natural-sounding sentences.
[1089] "Means of presenting generated sentence candidates to the user" refers to an interface that visually displays multiple generated sentences as candidates on a display or screen, allowing the user to select one.
[1090] "Means for the user to select generated sentence candidates" refers to the means by which the user selects the desired sentence from among multiple displayed sentence candidates.
[1091] "A means by which a user can give instructions to operate a robot using sign language" refers to an interface for transmitting operating instructions to a robot or machine through sign language actions.
[1092] A "robot operating means for executing instructions based on recognized sign language" is a control device that allows a robot to perform specific operations based on instructions recognized as sign language actions.
[1093] This invention relates to a system that captures a user's sign language movements and analyzes the resulting movement data. This system enables users to efficiently and intuitively generate text using sign language and to operate robots and give instructions in a factory setting.
[1094] hardware
[1095] The server and terminal include the following hardware:
[1096] Camera (e.g., Logitech C920): Captures the user's sign language movements in real time.
[1097] Robot control device: A device for controlling a robot based on recognized sign language movements.
[1098] software
[1099] The following software will be used:
[1100] Python: A programming language used for the overall system implementation.
[1101] OpenCV (CV2): A library for capturing and pre-processing camera footage.
[1102] Keras: A machine learning framework for running sign language gesture recognition models.
[1103] Natural language processing libraries (e.g., spaCy, GPT model): These are libraries for generating text candidates based on sign language actions.
[1104] System operation
[1105] The server receives video data of sign language movements transmitted from the user's camera. The video data is first preprocessed using OpenCV and then fed into a Keras model. The Keras model uses a pre-trained machine learning algorithm to convert the sign language movements into corresponding words.
[1106] The converted words are then transformed into multiple sentence candidates using a natural language processing library (e.g., spaCy or the GPT model). These sentence candidates are sent from the server to the terminal and displayed visually on the terminal's screen.
[1107] The user selects their desired sentence from several sentence options displayed on the screen. This selection is made through touch operations or finger movements. The final selected sentence is displayed on the device and then sent to the robot.
[1108] The robot performs specific actions based on the selected sentence. For example, if the user performs the sign language action for "pick up the part and move it," the robot will receive instructions such as "pick up the part and move it 5 meters."
[1109] Specific example
[1110] When a worker performs a sign language gesture for "pick up the part and move it," the server analyzes it and recognizes it as a corresponding word. The server then uses natural language processing technology to generate sentence options such as "pick up the part and move it 5 meters" or "take the part and move it to the designated location." These options are sent to the terminal, and the user selects the one sentence they prefer. Finally, the robot performs the action based on the selected instruction.
[1111] Example of a prompt
[1112] Example of sign language recognition results for use in analyzing sign language movements and generating text: Please generate phrases corresponding to the sign language words "part," "pick up," and "move." Possible instruction sentences would be something like, "Pick up the part and move 5 meters."
[1113] This invention enables intuitive and efficient operation and instruction using sign language in factories. Furthermore, it makes operation easier and improves work efficiency for workers who primarily use sign language for communication.
[1114] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1115] Step 1:
[1116] The server receives video data of sign language movements transmitted from the user's camera in real time. This video data is data captured by the camera of the sign language movements performed by the user. The data input is a video stream from the camera, and the output is the frame data of that video. Specifically, the camera continuously captures the user's hand movements and sends each frame to the server as digital data.
[1117] Step 2:
[1118] The server analyzes the received video data. OpenCV is used for preprocessing the video. In this step, the necessary sign language movements are extracted from the video, and noise is removed. The input for this process is frame data, and the output is preprocessed image data of the sign language movements. Specifically, this involves image smoothing and edge detection to enhance the shape of the hands.
[1119] Step 3:
[1120] The server supplies pre-processed video data to a Keras model, which runs a machine learning algorithm that converts sign language gestures into corresponding words. In this step, the input is pre-processed image data, and the output is the recognized words. Specifically, the model analyzes certain features in the image and maps them to pre-trained words.
[1121] Step 4:
[1122] The server generates multiple sentence candidates using a natural language processing library based on the recognized words. In this step, the input is the recognized words, and the output is a list of sentence candidates. Specifically, an NLP library (e.g., spaCy or GPT model) analyzes the word combinations and generates candidates as meaningful sentences.
[1123] Step 5:
[1124] The terminal presents the user with sentence suggestions sent from the server. In this step, the input is a list of generated sentence suggestions, and the output is the sentence displayed on the user's screen. Specifically, the terminal displays the sentence suggestions on its screen so that the user can visually confirm them.
[1125] Step 6:
[1126] The user selects a desired sentence from several sentence options displayed on the screen. In this step, the input is the sentence options displayed on the screen, and the output is the sentence selected by the user. Specifically, the user uses a touchscreen or mouse to select the desired sentence.
[1127] Step 7:
[1128] The terminal sends the text selected by the user to the server, which then distributes it to the robot control unit. In this step, the input is the text selected by the user, and the output is the control instruction for the robot. Specifically, the terminal sends the selected text to the server, and the server relays that instruction to the robot.
[1129] Step 8:
[1130] The robot performs specific operations based on the instructions it receives. In this step, the input is the instruction sent to the robot, and the output is the robot's physical movement. For example, based on the instruction "pick up and move the part," the robot arm picks up the part and moves it to the designated location.
[1131] This system enables intuitive robot operation using sign language, leading to increased efficiency in factory operations.
[1132] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1133] This invention combines a system that captures a user's sign language movements, analyzes the movement data to recognize specific words, with an emotion engine that recognizes the user's emotions. This system enables users to efficiently and intuitively generate sentences using sign language, and in addition, to generate sentences that include appropriate emotional expressions.
[1134] The user performs sign language actions through a device equipped with a camera and sensors. The device uses the camera to capture these sign language actions in real time. The camera continuously acquires video data and transmits it to a server.
[1135] The server analyzes the received video data and converts sign language movements into specific words. This uses image processing techniques and machine learning algorithms. Specifically, the server recognizes the shape and movement patterns of the hands and identifies the corresponding words based on a pre-trained model.
[1136] Next, the server generates sentence candidates based on the recognized words. Natural language processing (NLP) techniques are used here. The server generates multiple sentence candidates based on the recognized word combinations and lists them.
[1137] Furthermore, this system incorporates an emotion engine. The emotion engine analyzes the user's facial expressions and tone of voice to identify the user's emotional state. For example, if a user is smiling while performing sign language, the emotion engine recognizes that the user is "happy."
[1138] Based on the emotional state identified by the emotion engine, the server adjusts the tone and content of the generated text suggestions. For example, if the user is "happy," the server prioritizes generating text suggestions with a positive tone. Conversely, if the user is "sad," the server generates text suggestions that are more empathetic to that emotion.
[1139] The generated sentence candidates are sent to the terminal, which then visually presents them to the user. Specifically, the candidate sentences are displayed on the screen, allowing the user to select one. The user uses finger movements or touch gestures to select the desired sentence from the presented candidates.
[1140] The device detects the user's selection and displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others. The introduction of an emotion engine is expected to result in a more accurate and richer expression of the message.
[1141] For example, if a user performs the sign language actions for "I," "go," and "school," and the emotion engine recognizes the user's emotional state as "happy," the system will generate positive sentence options such as "I'm looking forward to going to school" or "I can't wait to go to school." The user then selects their desired sentence from these options, and the device displays the selected sentence as the final output.
[1142] In this way, the present invention is a system that can streamline communication using sign language and reflect the emotions of sign language users, thereby realizing more natural and enriching communication.
[1143] The following describes the processing flow.
[1144] Step 1:
[1145] The user performs sign language actions. A device equipped with a camera and sensors captures these sign language actions in real time. Specifically, the device uses its camera to continuously acquire video data, accurately capturing the shape and movement of the hand.
[1146] Step 2:
[1147] The terminal sends the acquired video data to the server. The server receives this video data and begins analysis. Data compression technology is used during data transfer for efficient transmission.
[1148] Step 3:
[1149] The server analyzes video data and converts sign language movements into specific words. The server uses image processing technology to analyze the contours and movements of the hands and maps the recognized movements to corresponding words through machine learning algorithms. Specifically, the server analyzes the shape, position, and movement patterns of the hands to recognize words such as "subject," "predicate," and "object."
[1150] Step 4:
[1151] The server uses natural language processing (NLP) techniques to generate sentence candidates based on the recognized words. For example, from the recognized words "I," "go," and "school," several sentences such as "I go to school" and "It is I who go to school" are created. The server then lists these sentence candidates.
[1152] Step 5:
[1153] The emotion engine analyzes the user's facial expressions and tone of voice to recognize their emotions. The device uses its camera and microphone to capture the user's facial expressions and voice data, and sends this data to the server.
[1154] Step 6:
[1155] The server analyzes the facial and voice data it receives to identify the user's emotional state. The emotion engine uses facial recognition and voice analysis technologies to recognize the user's emotional state, such as "happy," "sad," or "surprised."
[1156] Step 7:
[1157] The server adjusts the tone and content of the generated sentence suggestions based on the identified emotional state. For example, if the user is perceived as "happy," the server prioritizes generating sentence suggestions with a positive tone. Specifically, it might generate sentence suggestions that include emotional expressions such as, "I'm looking forward to going to school."
[1158] Step 8:
[1159] The server sends generated sentiment-reflecting sentence candidates to the terminal, which then visually presents them to the user. Multiple candidate sentences are displayed on the screen in a list format, with each candidate numbered or highlighted.
[1160] Step 9:
[1161] The user selects their desired text from a list of suggested texts. The user makes the selection using finger movements or touch gestures. The device captures the user's selection actions and sends that information to the server.
[1162] Step 10:
[1163] The server receives the user's selection information and finally confirms the selected text. The selected text is then sent from the server to the terminal.
[1164] Step 11:
[1165] The device visually displays the final selected text. This allows the user to confirm what they want to convey and communicate it to others.
[1166] Through the processing steps described above, the present invention realizes efficient and intuitive text generation and communication using sign language. Furthermore, the introduction of an emotion engine enables natural and rich communication that also reflects emotional expression.
[1167] (Example 2)
[1168] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1169] Conventional sign language recognition systems only recognize sign language movements, and the generated text does not reflect the user's emotions, resulting in the problem of inaccurately conveying the intended message. Furthermore, providing a diverse range of example sentences during text generation is difficult, limiting communication possibilities.
[1170] The identification processing by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words, means for presenting the generated sentence candidates to the user, means for the user to select a generated sentence candidate, and means for recognizing the user's emotional state and adjusting the tone and content of the sentence candidates based on that emotional state. This makes it possible to generate accurate sentences based on the user's sign language actions and to provide a variety of sentence candidates that reflect the user's emotions.
[1171] A "user" is a person who uses this system and is the entity that inputs data through sign language actions.
[1172] "Means for capturing movements" refers to devices or technologies that use cameras or sensors to acquire a user's sign language movements in real time.
[1173] "Motion data" refers to video information and sensor data that digitally records the user's sign language movements.
[1174] "Means of analysis" refers to methods and techniques for analyzing acquired motion data using machine learning algorithms and image processing technologies, and converting it into specific words.
[1175] A "specific word" is a string of characters with meaning recognized from the analyzed motion data, representing a conversion of sign language movements into language.
[1176] "Methods for generating sentence candidates" refer to methods and techniques for creating multiple sentences based on recognized words using natural language processing technology or generative AI models.
[1177] "Means for presenting generated sentence candidates to the user" refers to a technology that visually displays multiple sentence candidates to the user using a display device such as a screen.
[1178] "Means for users to select generated text candidates" refers to a device or technology that allows a user to select a desired text from a visually displayed list of text candidates using a touchscreen or finger movements.
[1179] "Means of recognizing emotional states" refers to methods and technologies that use facial recognition technology or voice analysis technology to identify a user's emotions.
[1180] "Means of adjusting tone and content" refers to methods and techniques for appropriately modifying the style and content of candidate texts generated based on the recognized emotional state of the user.
[1181] This invention relates to a system that captures a user's sign language movements, analyzes the movement data to recognize specific words, and further recognizes the user's emotions to generate sentence candidates based on those emotions. This system efficiently and intuitively generates sentences based on the user's sign language movements and reflects emotions, enabling natural communication.
[1182] hardware
[1183] The system primarily uses the following hardware:
[1184] 1. Device: An electronic device equipped with a camera or sensors (e.g., smartphone, tablet, etc.).
[1185] 2. Camera: A high-resolution camera that captures sign language movements in real time.
[1186] software
[1187] The system uses the following software and technologies:
[1188] 1. Image processing technology: Use OpenCV or similar tools to analyze sign language movements.
[1189] 2. Machine learning algorithm: TensorFlow or PyTorch is used to recognize specific words from behavioral data.
[1190] 3. Natural Language Processing (NLP) Technology: Generative AI models such as GPT-3 are used to generate sentence candidates based on recognized words.
[1191] 4. Emotion Recognition Technology: Use Amazon Rekognition or Azure Face API to identify the user's emotional state and adjust suggested text accordingly.
[1192] Data processing and calculation
[1193] Capture and transmit sign language movements
[1194] The user performs sign language gestures in front of the device. The device uses its built-in camera to capture the sign language gestures in real time and sends the video data frame by frame to the server. HTTP or WebSocket protocol is used for transmission.
[1195] Video data analysis and word recognition
[1196] The server analyzes the received video data using image processing technologies such as OpenCV and TensorFlow, as well as machine learning algorithms, to recognize specific words. It detects hand shapes and movement patterns and converts them into corresponding words based on a pre-trained model.
[1197] Sentence candidate generation
[1198] The server generates sentence candidates using a generative AI model such as GPT-3 based on recognized words. It takes a prompt message as input and generates and lists multiple sentence candidates. For example, if the user inputs the words "I," "go," and "school," the server will use a prompt message like the following:
[1199] Sign language words to input: I, go, school
[1200] Emotional state: Happy
[1201] Generation example:
[1202] I look forward to going to school.
[1203] I can't wait to go to school.
[1204] Identifying and adjusting emotional states
[1205] The server uses emotion recognition technologies such as Amazon Rekognition and Azure Face API to analyze the user's facial expressions and voice tone from video data and identify their emotional state. For example, if a user is smiling while using sign language, the server recognizes this as "happy." Based on the identified emotional state, the server adjusts the tone and content of the generated text suggestions.
[1206] Presentation and selection of sentence options
[1207] The generated sentence suggestions are sent to the device, which displays them on its screen. The user selects the desired sentence using touch gestures or finger movements. For example, "I'm looking forward to going to school" and "I can't wait to go to school" might be displayed, and the user selects from among them.
[1208] Display of final output
[1209] The device ultimately displays the text selected by the user. This allows the user to accurately convey what they want to communicate, including their emotions, to others.
[1210] As described above, the system of the present invention enables natural communication that reflects the user's sign language movements and emotions.
[1211] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1212] Step 1: The user performs a sign language action.
[1213] The user performs sign language actions in front of the device's camera. For example, they might perform sign language actions for "I," "go," and "school." The user's sign language actions are the input and are captured as video data.
[1214] Step 2: The device captures sign language movements and sends the video data to the server.
[1215] The device uses a built-in high-resolution camera to capture the user's sign language movements in real time. The captured video data is acquired frame by frame and transmitted to the server via Wi-Fi or mobile data communication. The input is video data of the sign language movements, and the output is video data transferred to the server.
[1216] Step 3: The server analyzes the video data and converts sign language actions into words.
[1217] The server uses image processing techniques such as OpenCV and TensorFlow / PyTorch, along with machine learning algorithms, to analyze the received video data. It detects hand shapes and movement patterns and uses a pre-trained model to convert those movements into corresponding specific words. The input is video data, and the output is a list of recognized words.
[1218] Step 4: The server generates sentence suggestions based on the words.
[1219] The server generates sentence candidates using a generative AI model (e.g., GPT-3) based on recognized word combinations. A prompt sentence is entered, and multiple sentence candidates are generated and listed. The input consists of a list of recognized words and a prompt sentence, and the output is a list of sentence candidates. For example, if the words "I," "go," and "school" are entered, the server generates sentence candidates such as "I look forward to going to school" and "I can't wait to go to school."
[1220] Step 5: The server uses the emotion engine to identify the emotional state and adjust the suggested sentences.
[1221] The server uses emotion recognition technology (such as Amazon Rekognition or Azure Face API) to analyze facial expressions and voice tone to identify the user's emotional state. For example, if the user is smiling and using sign language in video data, it will recognize that the user is "happy." The input is the user's video data, and the output is the identified emotional state. Subsequently, the tone and content of the generated sentence candidates are adjusted according to the emotional state. For example, if the user is happy, positive sentences will be prioritized.
[1222] Step 6: The server sends the adjusted text candidates to the terminal.
[1223] The server sends the adjusted sentence candidates to the terminal in JSON format. The input is the adjusted sentence candidates, and the output is the data sent to the terminal. Specifically, sentence candidates such as "I'm looking forward to going to school" and "I can't wait to go to school" are sent.
[1224] Step 7: The device presents text suggestions to the user, and the user makes a selection.
[1225] The device displays received text suggestions on its screen. The user selects the desired text using the touchscreen or finger movements. Text suggestions are sent to the device as input, and the text selected by the user is returned as output. For example, this might involve displaying multiple sentences on the screen and the user selecting "I am looking forward to going to school."
[1226] Step 8: The device displays the final text.
[1227] The device ultimately displays the text selected by the user. The user selects the text as input, and the final displayed text is the output. For example, the text "I am looking forward to going to school" might be displayed on the screen. This display allows the user to accurately convey what they want to communicate to others.
[1228] (Application Example 2)
[1229] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1230] This invention aims to improve the efficiency and accuracy of communication in the workplace by converting sign language used by hearing-impaired workers into text in real time, while also considering the worker's emotional state during the text generation. Furthermore, it aims to reduce stress and anxiety experienced by workers by providing appropriate feedback and support that reflects their emotional state, thereby improving work safety and productivity.
[1231] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1232] In this invention, the server includes means for capturing user actions, means for analyzing the captured action data and recognizing specific words, means for generating sentence candidates based on the recognized words and the user's emotional state, means for presenting the generated sentence candidates to the user, and means for the user to select a generated sentence candidate. This makes sign language communication more efficient and reflects the user's emotional state, enabling richer and more accurate communication.
[1233] "Means for capturing user actions" refers to devices that use cameras or sensors to acquire the physical movements performed by a user as digital data.
[1234] "Means for analyzing captured motion data and recognizing specific words" refers to a device or software that uses machine learning algorithms or image processing techniques to identify specific words or meanings based on user motion data.
[1235] "Means for generating sentence candidates based on recognized words and the user's emotional state" refers to a device or software that generates multiple sentence candidates using words recognized from the user's sign language actions and the user's emotional state analyzed by an emotion engine.
[1236] "Means for presenting generated sentence options to the user" refers to a display that visually shows the user the options for the generated sentences, or a device that provides audio notifications.
[1237] "Means for the user to select generated sentence candidates" refers to an interface for the user to select a desired sentence from the presented sentence candidates, and includes touchscreens and other input devices.
[1238] "Sign language" refers to a means of communication using hand and finger movements, which are physical actions used to convey specific words or meanings.
[1239] "Emotional state" refers to the user's emotional and mood state, and is identified by analyzing facial expressions, tone of voice, and other factors.
[1240] "Natural language processing technology" refers to technologies that enable computers to understand and generate human language, and includes algorithms for text analysis and generation.
[1241] "Emotional analysis technology" is a technology that analyzes a user's facial expressions, voice tone, body movements, etc., to identify the user's current emotional state.
[1242] The embodiments for carrying out this invention will be described in detail below.
[1243] First, the system program is generated. This program captures the user's sign language movements, analyzes the movement data, and recognizes specific words. It also incorporates an emotion engine that recognizes the user's emotional state, and includes emotional expressions in the generated sentences.
[1244] The system consists of the following main hardware and software components:
[1245] 1. Means for capturing user actions
[1246] Cameras and sensors are used to capture the user's hand and finger movements in real time. Specifically, webcams and dedicated motion capture devices are used.
[1247] 2. A means of analyzing captured motion data and recognizing specific words.
[1248] Machine learning algorithms and image processing techniques are used to analyze captured sign language movements. Specifically, libraries such as TensorFlow and OpenCV are used.
[1249] 3. A means of generating sentence candidates based on recognized words and the user's emotional state.
[1250] This system generates multiple text candidates using natural language processing (NLP) and emotion analysis techniques. Specifically, it employs an emotion engine (e.g., EmotionEngine) and an NLP model (e.g., NLPModel).
[1251] 4. Means for presenting generated text candidates to the user
[1252] The system uses a display to present the user with generated text suggestions. This can be done using a computer monitor or projector.
[1253] 5. Means for the user to select generated text suggestions
[1254] The user selects the desired text using a touchscreen or other interface. Specifically, touch panels and pointing devices are used.
[1255] Furthermore, specific examples are given below.
[1256] As an example, consider a case where a hearing-impaired factory worker communicates in sign language that "the machine is broken." First, a camera captures the sign language movements, and the system analyzes the movements to recognize that "the machine is broken." Simultaneously, an emotion engine analyzes the worker's facial expressions to read their feelings of anxiety. Based on this, the system generates a suggested sentence, "The machine is broken. I am very anxious," and displays it on the screen. The worker can then select this sentence using a touch panel to inform other workers or supervisors.
[1257] Examples of prompt statements include the following:
[1258] "Imagine a system that uses sign language to report the situation within a factory. For example, if a worker expresses in sign language that 'the machine is broken' and shows an anxious expression, build a system that analyzes the sign language and emotions to provide appropriate feedback. This system combines sign language recognition and emotion recognition technologies to generate a detailed report."
[1259] Thus, this invention streamlines communication using sign language and reflects the emotions of the workers, thereby achieving more natural and enriching communication.
[1260] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1261] Step 1:
[1262] Camera-based motion capture
[1263] The server uses a camera to capture the user's sign language movements in real time. The input is the user's hand and finger movements, and the output is video frame data. This frame data is then sent to the next analysis step.
[1264] Step 2:
[1265] Sign language motion recognition
[1266] The server analyzes captured video frame data using image processing techniques and machine learning algorithms to recognize specific words. At this stage, TensorFlow and OpenCV are used to analyze hand shapes and movement patterns. The input is video frame data, and the output is the recognized word.
[1267] Step 3:
[1268] Emotion analysis
[1269] The server analyzes the user's facial expressions and voice tone from the captured video frame data and uses an EmotionEngine to identify the user's emotional state. The input is the same video frame data, and the output is the analyzed emotional state.
[1270] Step 4:
[1271] Sentence generation
[1272] The server generates multiple sentence candidates using natural language processing (NLP) techniques based on recognized words and analyzed sentiment states. An NLP model is used at this stage. The input is recognized words and sentiment states, and the output is multiple sentence candidates.
[1273] Step 5:
[1274] Suggestion of sentence candidates
[1275] The terminal visually displays the generated sentence candidates on its screen and presents them to the user. The input consists of multiple sentence candidates, and the output includes the sentence candidates displayed on the screen. The user is then ready to make a selection.
[1276] Step 6:
[1277] Selection of sentence candidates
[1278] The user selects their desired sentence from the presented options. The device acquires the user's selection data using a touchscreen or pointing device. The input is the user's selection operation, and the output is the final sentence selected by the user.
[1279] Step 7:
[1280] Output of selected text
[1281] The device displays the final text selected by the user on its screen, confirming the intent of the communication. The input is the text selected by the user, and the output is the text that is finally displayed on the device. This ensures that the user's intent is conveyed to the other person.
[1282] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1283] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1284] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1285] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1286] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1287] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1288] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1289] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1290] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1291] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1292] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1293] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1294] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1295] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1296] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1297] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1298] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1299] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1300] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1301] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1302] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[1303] The following is further disclosed regarding the embodiments described above.
[1304] (Claim 1)
[1305] Means for capturing user actions,
[1306] A means of analyzing captured motion data and recognizing specific words,
[1307] A means of generating sentence candidates based on recognized words,
[1308] A means of presenting generated text suggestions to the user,
[1309] A system including means for the user to select generated text candidates.
[1310] (Claim 2)
[1311] The system according to claim 1, wherein the captured motion data is a sign language movement of a user using a camera, and includes an analysis means for identifying said sign language movement.
[1312] (Claim 3)
[1313] The system according to claim 1, wherein the means for generating text candidates is to create multiple text candidates using natural language processing technology.
[1314] "Example 1"
[1315] (Claim 1)
[1316] Means for capturing user actions,
[1317] A means of analyzing captured motion data and recognizing specific words,
[1318] A means of generating sentence candidates based on recognized words,
[1319] A means of presenting generated text suggestions to the user,
[1320] A means for the user to select generated text suggestions,
[1321] A means of displaying the selected text,
[1322] Methods of using cameras and sensors for capture,
[1323] A method using image processing techniques to identify the shape and movement patterns of the hand,
[1324] A means of recognizing words using a trained machine learning algorithm,
[1325] A means for generating text candidates using natural language processing technology,
[1326] A means for the device to send video data acquired by the camera to a server,
[1327] A system that includes this.
[1328] (Claim 2)
[1329] The system according to claim 1, wherein the captured motion data is a sign language movement of a user using a camera, and includes an analysis means for identifying said sign language movement.
[1330] (Claim 3)
[1331] The system according to claim 1, wherein the means for generating text candidates is to create multiple text candidates using natural language processing technology.
[1332] "Application Example 1"
[1333] (Claim 1)
[1334] Means for capturing user actions,
[1335] A means of analyzing captured motion data and recognizing specific words,
[1336] A means of generating sentence candidates based on recognized words,
[1337] A means of presenting generated text suggestions to the user,
[1338] A means for the user to select generated text suggestions,
[1339] A means for users to give instructions to operate a robot using sign language,
[1340] A robot operating means that executes instructions given in recognized sign language,
[1341] A system that includes this.
[1342] (Claim 2)
[1343] The system according to claim 1, wherein the captured motion data is a sign language movement of a user using a camera, and includes an analysis means for identifying said sign language movement.
[1344] (Claim 3)
[1345] The system according to claim 1, wherein the means for generating text candidates is to create multiple text candidates using natural language processing technology.
[1346] "Example 2 of combining an emotion engine"
[1347] (Claim 1)
[1348] Means for capturing user actions,
[1349] A means of analyzing captured motion data and recognizing specific words,
[1350] A means of generating sentence candidates based on recognized words,
[1351] A means of presenting generated text suggestions to the user,
[1352] A means for the user to select generated text suggestions,
[1353] A means for recognizing the user's emotional state and adjusting the tone and content of suggested texts based on that emotional state,
[1354] A system that includes this.
[1355] (Claim 2)
[1356] The system according to claim 1, wherein the captured motion data is a sign language movement of a user using a camera, and includes an analysis means for identifying said sign language movement.
[1357] (Claim 3)
[1358] The system according to claim 1, wherein the means for generating text candidates involves creating multiple text candidates using natural language processing technology, and further generates text candidates based on a prompt sentence using a generation AI model.
[1359] "Application example 2 when combining with an emotional engine"
[1360] (Claim 1)
[1361] Means for capturing user actions,
[1362] A means of analyzing captured motion data and recognizing specific words,
[1363] A means for generating sentence candidates based on recognized words and the user's emotional state,
[1364] A means of presenting generated text suggestions to the user,
[1365] A system including means for the user to select generated text candidates.
[1366] (Claim 2)
[1367] The system according to claim 1, wherein the captured motion data is a user's sign language movements using a camera, and the system includes an analysis means for identifying the sign language movements and the user's emotional state.
[1368] (Claim 3)
[1369] The system according to claim 1, wherein the means for generating text candidates is to create multiple text candidates using natural language processing technology and emotion analysis technology. [Explanation of Symbols]
[1370] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for capturing user actions, A means of analyzing captured motion data and recognizing specific words, A means of generating sentence candidates based on recognized words, A means of presenting generated text suggestions to the user, A system including means for the user to select generated text candidates.
2. The system according to claim 1, wherein the captured motion data is a sign language movement of a user using a camera, and includes an analysis means for identifying said sign language movement.
3. The system according to claim 1, wherein the means for generating text candidates is to create multiple text candidates using natural language processing technology.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A