system

The system addresses real-time finger movement recognition challenges by using a camera and machine learning to process gestures, enhancing user interaction through intuitive device operation.

JP2026064728APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Conventional systems face challenges in accurately recognizing user finger movements in real time, leading to delays and misrecognition in image processing and gesture recognition, which can degrade user experience.

Method used

A system comprising a camera, image processing, feature point extraction, server, gesture recognition, and terminal that captures and processes finger movements in real time, using preprocessing techniques like grayscale conversion and noise reduction, and machine learning models to recognize gestures and generate commands.

Benefits of technology

Enables intuitive and efficient operation of devices using hand gestures, improving user experience by allowing touchless and convenient interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064728000001_ABST
    Figure 2026064728000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A camera means for acquiring the user's hand movements in real time, Image processing means for preprocessing image data acquired from the camera means, A means for extracting characteristic points of fingers from image data processed by the aforementioned image processing means, A server means for analyzing the aforementioned feature point data, A means for recognizing gestures based on feature point data analyzed by the server means, Means for interpreting and generating a specific command based on the gesture recognition means, A terminal means that executes the aforementioned command and provides feedback of the result to the user, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a conventional system, there are problems that it is difficult to accurately recognize the movement of a user's finger and perform an appropriate operation in real time, or a delay occurs. Also, in preprocessing of images and gesture recognition, there are problems such as misrecognition and reaction delay due to a lack of a highly accurate and efficient algorithm. Furthermore, if the feedback to a user's operation is not prompt, the user experience may decline.

Means for Solving the Problems

[0005] This invention provides a system including a camera, image processing, feature point extraction, server, gesture recognition, command interpretation / generation, and terminal for acquiring the user's finger movements in real time. The camera captures the user's finger movements in real time, and the image processing performs preprocessing on the collected image data, such as grayscale conversion, noise reduction, and contrast adjustment. The feature point extraction extracts feature points of the fingers from the preprocessed data, and the server analyzes the feature point data using a machine learning model to recognize the type of gesture. The gesture recognition interprets and generates a specific command based on the recognized gesture. The terminal executes the command and provides feedback to the user.

[0006] A "camera device" is a device used to capture the user's hand movements in real time.

[0007] "Image processing means" refers to a device or algorithm that performs preprocessing on image data acquired from a camera means, such as grayscale conversion, noise reduction, and contrast adjustment.

[0008] "Feature point extraction means" refers to a device or algorithm that extracts feature points of fingers from image data processed by image processing means.

[0009] A "server system" is a computer system that receives feature point data and analyzes it using a machine learning model.

[0010] "Gesture recognition means" refers to a device or algorithm that recognizes the type of gesture based on feature point data analyzed by the server means.

[0011] "Command interpretation and generation means" refers to a device or algorithm that interprets and generates a specific command based on gesture recognition means.

[0012] A "terminal means" is a device that executes commands generated by a command interpretation / generation means and provides feedback of the results to the user. [Brief explanation of the drawing]

[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiments for Carrying Out the Invention

[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0015] First, the terms used in the following description will be explained.

[0016] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0017] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0021] [First Embodiment]

[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0034] This invention is a system for recognizing the movements of a user's fingers in real time and performing operations based on those movements. This system consists of the following components: a camera, an image processing unit, a feature point extraction unit, a server, a gesture recognition unit, a command interpretation and generation unit, and a terminal unit.

[0035] System Configuration

[0036] Camera means

[0037] The camera built into the device captures the user's hand movements in real time. The captured image data is processed in real time.

[0038] Image processing means

[0039] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0040] Feature point extraction means

[0041] Feature points, such as the joints and fingertips of the fingers, are extracted from the preprocessed image data. This feature point extraction uses a specific algorithm to convert the shape and position of the fingers into numerical data.

[0042] Server Means

[0043] After receiving feature point data transmitted from the terminal, the server performs analysis using a machine learning model. This machine learning model has been pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[0044] Gesture recognition means

[0045] The server classifies the type of gesture based on the feature point data it analyzes. Based on this classification, it interprets and generates the relevant commands.

[0046] Command interpretation and generation means

[0047] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[0048] Terminal means

[0049] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback."

[0050] Specific example

[0051] When playing music using gestures

[0052] The user makes a "play" gesture (for example, drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera and captures the user's hand movements in real time. The acquired image is pre-processed by the device, including grayscale conversion, noise reduction, and contrast adjustment. Feature points of the hands are extracted from the pre-processed image, and this data is sent to the server.

[0053] The server uses a machine learning model to analyze feature point data and recognizes it as a "play" gesture. Based on this, the server generates a music playback command and sends it to the terminal. The terminal receives this command, launches the music player application, starts playing music, and provides feedback to the user saying, "Music playback will begin."

[0054] This system allows users to intuitively perform various operations using only hand gestures, providing high operability and a superior user experience.

[0055] The following describes the processing flow.

[0056] Step 1:

[0057] The device activates its built-in camera. The camera begins capturing the user's hand movements in real time.

[0058] Step 2:

[0059] The device acquires captured images frame by frame in real time. This allows for the continuous collection of the user's hand movements.

[0060] Step 3:

[0061] The device performs preprocessing on the acquired image data. Specifically, it converts the image to grayscale, applies a noise reduction filter, and adjusts the contrast. This improves the accuracy of subsequent processing.

[0062] Step 4:

[0063] The device segments the finger area from the pre-processed image. This is done using color distribution and contour detection algorithms.

[0064] Step 5:

[0065] The device extracts feature points from segmented finger regions. Specifically, it identifies the positions of finger joints and fingertips and records them as numerical data.

[0066] Step 6:

[0067] The terminal sends the extracted feature point data to the server. This transmission is done via the internet or a local network.

[0068] Step 7:

[0069] The server receives feature point data and prepares it for analysis. The received data is then loaded into the system.

[0070] Step 8:

[0071] The server analyzes the data using machine learning models. This involves algorithms such as neural networks and support vector machines.

[0072] Step 9:

[0073] The server recognizes specific gestures based on data analysis. For example, drawing a circle with the thumb and index finger is recognized as the "play" gesture.

[0074] Step 10:

[0075] The server interprets and generates a specific command (e.g., play music) based on the recognized gesture.

[0076] Step 11:

[0077] The server sends a feedback message (e.g., "Starting music playback") to the terminal along with the generated command.

[0078] Step 12:

[0079] The device displays a feedback message to the user. For example, it might display "Starting music playback" on the screen.

[0080] Step 13:

[0081] The terminal executes a command. Specifically, it launches the music player application and starts playing music.

[0082] (Example 1)

[0083] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0084] Currently, many users utilize smart devices to access a variety of applications and services. However, these operations are typically performed using touchscreens or physical buttons, often resulting in cumbersome user experiences. Furthermore, traditional methods make intuitive operation and the use of hand gestures difficult, making them particularly inconvenient when hands are dirty or when performing other tasks. A system is needed to solve these problems and provide a more intuitive and convenient way to operate the device.

[0085] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0086] In this invention, the server includes: a shooting device means for acquiring the user's finger movements in real time; an image conversion means for preprocessing image data acquired from the shooting device means; a means for extracting finger feature points from the image data processed by the image conversion means; an analysis device means for analyzing the feature point data; a means for recognizing gestures based on the feature point data analyzed by the analysis device means; a means for interpreting and generating specific commands based on the gesture recognition means; and a terminal device means for executing the commands and feeding the results back to the user. This allows the user to operate intuitively using finger movements, enabling touchless and convenient operation even while working.

[0087] "Photography device means" refers to a device for acquiring the movements of the user's fingers in real time, and includes cameras and the like.

[0088] The "image conversion means" is a means for preprocessing image data acquired from the imaging device means, and performs processing such as grayscale conversion, noise reduction, and contrast adjustment.

[0089] "Methods for extracting feature points of fingers" refer to methods for extracting feature points such as the joints and fingertips of fingers from preprocessed image data, and primarily utilize algorithms and deep learning models.

[0090] "Analysis device means" refers to a means of receiving feature point data transmitted from a terminal and performing analysis using a machine learning model.

[0091] "Means for recognizing gestures" refers to means for classifying and recognizing the type of gesture based on feature point data analyzed by an analysis device.

[0092] "Means for interpreting and generating commands" refers to a method for generating specific commands based on recognized gestures and for interpreting those commands.

[0093] A "terminal device means" is a means of receiving commands sent from a server, actually executing those commands, and providing feedback of the results to the user.

[0094] This invention is a system for recognizing the movements of a user's fingers in real time and performing operations based on those movements. This system consists of the following components: a shooting device, an image conversion device, a device for extracting characteristic points of the fingers, an analysis device, a gesture recognition device, a command interpretation and generation device, and a terminal device.

[0095] System Configuration

[0096] Photography device means

[0097] The camera built into the device captures the user's hand movements in real time. The captured image data is processed in real time.

[0098] Image conversion means

[0099] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. For example, the OpenCV library is used for this preprocessing. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0100] A method for extracting characteristic points of the fingers.

[0101] Feature points, such as the joints and fingertips of the fingers, are extracted from the preprocessed image data. For this feature point extraction, a model like the MediaPipe Hand is used. The feature points are obtained as numerical data representing the shape and position of the fingers.

[0102] analysis equipment means

[0103] After receiving feature point data transmitted from the terminal, the server performs analysis using a machine learning model. This machine learning model is, for example, a gesture recognition model trained with TENSORFLOW®. The server uses this model to recognize the type of gesture based on the newly received data.

[0104] Means of recognizing gestures

[0105] The server classifies the gesture type based on the analysis results. Based on this classification, it interprets and generates the relevant commands.

[0106] Means for interpreting and generating commands

[0107] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[0108] Terminal device means

[0109] The device receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, the device launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback."

[0110] Specific example

[0111] When playing music using gestures

[0112] The user makes a "play" gesture (for example, drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera and captures the user's hand movements in real time. The device preprocesses the acquired image by performing grayscale conversion, noise reduction, contrast adjustment, etc. Feature points of the hands are extracted from the preprocessed image and this data is sent to the server. The server analyzes the feature point data using a machine learning model and recognizes that it is a "play" gesture. Based on this result, the server generates a music playback command and sends it to the device. The device receives this command, launches the music player application, starts music playback, and provides feedback to the user saying, "Music playback will begin."

[0113] This system allows users to intuitively perform various operations using only hand gestures. An example of a prompt to be input to the AI ​​model is as follows:

[0114] Example of a prompt

[0115] Prompt: Create a system that recognizes a user's finger gesture for "play" and plays music accordingly. The steps are as follows:

[0116] 1. Use the device's camera to record the movement of your fingers.

[0117] 2. Preprocess the acquired image data (grayscale conversion, noise reduction, contrast adjustment).

[0118] 3. Extract feature points from the preprocessed image.

[0119] 4. Send the extracted feature point data to the server.

[0120] 5. The server analyzes the feature point data and recognizes the gesture.

[0121] 6. Generate a command based on the recognition result and send it to the terminal.

[0122] 7. The terminal receives the command and plays the music.

[0123] This allows users to operate the system with high operability and an intuitive interface.

[0124] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0125] Step 1:

[0126] Camera activation and image capture

[0127] The device activates its built-in camera and captures the user's finger movements in real time. Suppose the user performs a "play" gesture in front of the camera (for example, drawing a circle with their thumb and index finger). The input is the user's finger movements, and the output is real-time image data. This image data is retained for subsequent processing.

[0128] Specific actions:

[0129] The user moves their fingers in front of the device, and the device activates its camera to take a picture.

[0130] Step 2:

[0131] Image preprocessing

[0132] The device performs preprocessing on image data acquired in real time, including grayscale conversion, noise reduction, and contrast adjustment. This improves the quality of the image data and facilitates subsequent processing. The input is the image data from step 1, and the output is the preprocessed image data.

[0133] Specific actions:

[0134] The device converts the image data to grayscale, applies a noise reduction filter, and adjusts the image contrast.

[0135] Step 3:

[0136] Feature point extraction

[0137] The device extracts feature points of the fingers from pre-processed image data. For feature point extraction, for example, the MediaPipe Hand model is used. This model converts the positions of finger joints and fingertips into numerical data. Pre-processed image data is taken as input, and feature point data of the fingers is generated as output.

[0138] Specific actions:

[0139] The device uses the MediaPipe Hand model to calculate the positions of the joints and fingertips of the fingers and obtains numerical data indicating these positions.

[0140] Step 4:

[0141] Sending feature point data to the server

[0142] The device sends feature point data of the fingers to the server. This data is sent to the server via the network as a POST request. The input is feature point data, and the output is a request sent to the server.

[0143] Specific actions:

[0144] The terminal converts the feature point data into JSON format and sends a POST request to the server's API endpoint.

[0145] Step 5:

[0146] Gesture recognition and classification

[0147] The server uses the received feature point data to analyze it with a machine learning model and recognize the type of gesture. The model used is, for example, a gesture recognition model trained with TensorFlow. The input is feature point data, and the output is the gesture recognition result.

[0148] Specific actions:

[0149] The server inputs the received feature point data into a TensorFlow model and performs the analysis. The model recognizes the "play" gesture and generates the recognition result.

[0150] Step 6:

[0151] Command generation and sending

[0152] The server generates a command corresponding to the recognized gesture and sends that command to the terminal. This command could include, for example, an instruction to play music. The input is the gesture recognition result, and the output is the generated command.

[0153] Specific actions:

[0154] The server analyzes the gesture recognition results and generates a corresponding music playback command. The generated command is then sent to the terminal in JSON format.

[0155] Step 7:

[0156] Command execution and feedback

[0157] The terminal receives and executes a command sent from the server. It launches a music player application, starts playing music, and provides feedback to the user saying, "Starting music playback." The input is a command, and the output is music playback and a feedback message.

[0158] Specific actions:

[0159] The terminal analyzes the music playback command, launches the music player, and starts playback. At the same time, a message saying "Music playback will begin" is displayed on the screen to the user.

[0160] (Application Example 1)

[0161] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0162] In modern factories, improving work efficiency and ensuring worker safety are critical challenges. However, conventional robot operating systems often lack intuitive and rapid response capabilities, potentially leading to operational errors and increased worker burden. In particular, the presence of numerous buttons and switches increases the complexity of manual operation and makes operational errors more likely. To address these challenges, a system is needed that more intuitively recognizes worker movements and provides accurate operational feedback.

[0163] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0164] In this invention, the server includes a camera means for acquiring the user's finger movements in real time, an image processing means for preprocessing image data acquired from the camera means, a means for extracting finger feature points from the image data processed by the image processing means, a server means for analyzing the feature point data, a means for recognizing gestures based on the feature point data analyzed by the server means, a means for interpreting and generating specific commands based on the gesture recognition means, a terminal means for executing the commands and feeding the results back to the user, a means for operating the robot in a specific work environment, and a means for starting, stopping, or adjusting work based on gestures using the means. This enables the operator to intuitively and quickly operate the robot using only finger movements, thereby improving work efficiency and ensuring safety.

[0165] A "camera device" is a device used to capture the movements of the user's fingers in real time, and its role is to acquire image data.

[0166] "Image processing means" refers to techniques used to preprocess image data acquired from a camera and improve the quality of the processed image data.

[0167] The "feature point extraction means" is an algorithm for extracting feature points of fingers from pre-processed image data and converting their position and shape into numerical data.

[0168] A "server system" is a computer server that receives feature point data transmitted from a terminal and analyzes the data using a machine learning model.

[0169] "Gesture recognition means" refers to a technology that classifies and recognizes the type of user gesture based on feature point data analyzed by the server means.

[0170] The "command interpretation and generation means" is a means for interpreting and generating a corresponding command in response to a gesture recognized by the gesture recognition means.

[0171] "Terminal means" refers to equipment or devices that execute commands sent from server means and provide feedback of the results to the user.

[0172] "Means for operating robots in a specific work environment" refers to means used to direct specific actions or tasks when operating robots in a factory or work environment.

[0173] "Means for starting, stopping, or adjusting work based on gestures" refers to technology that recognizes the user's hand gestures and starts, stops, or adjusts the robot's movements in response to those gestures.

[0174] This invention is a system for intuitively operating a robot using specific gestures in a factory work environment. Detailed embodiments of this system are described below.

[0175] Key components of the system

[0176] Camera means

[0177] A camera system is used to capture the user's hand movements in real time. The camera, using a webcam or a dedicated high-resolution camera, captures the hand movements and acquires the data.

[0178] Image processing means

[0179] The acquired image data is preprocessed by image processing equipment. Specifically, grayscale conversion, noise reduction, and contrast adjustment are performed using image processing libraries such as OpenCV. This improves the quality of the image data and facilitates subsequent processing by feature point extraction equipment.

[0180] Feature point extraction means

[0181] Feature points, such as the joints and fingertip positions of the fingers, are extracted from the preprocessed image data. Specific algorithms (e.g., HOG or SIFT) are used to convert the shape and position of the fingers into numerical data.

[0182] Server Means

[0183] Feature point data is sent from the terminal to the server. The server analyzes the feature point data using a machine learning model. This machine learning model is built using frameworks such as TensorFlow and trained on a large amount of gesture data.

[0184] Gesture recognition means

[0185] The server classifies the type of gesture based on the analyzed feature point data. For example, if a user makes a "start" gesture, the server will recognize that gesture correctly.

[0186] Command interpretation and generation means

[0187] Based on the gesture recognition results, the server generates corresponding commands. Specifically, commands such as "start," "stop," and "adjust" are generated.

[0188] Terminal means

[0189] The generated command is sent to the terminal, which then executes the command. For example, if the command "start" is received, the robot will begin operation. The terminal also provides feedback to the user, such as "The robot will begin its work."

[0190] System operation example

[0191] For example, recognizing the "start" gesture will cause the conveyor belt in the factory to begin moving. Recognizing the "stop" gesture will cause the conveyor belt to stop. Recognizing the "adjust" gesture will adjust the robot's speed and direction. In this way, users can intuitively and quickly operate robots using only finger movements.

[0192] Example of a prompt

[0193] The following are examples of prompts to input into a generative AI model in a real-world usage scenario:

[0194] When the "start" gesture is recognized, the conveyor belts inside the factory begin to move.

[0195] When the "stop" gesture is recognized, the conveyor belt will stop.

[0196] When it recognizes the "adjust" gesture, it adjusts the robot's movement speed and direction.

[0197] As described above, the present invention improves work efficiency and worker safety within factories.

[0198] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0199] Step 1:

[0200] The terminal activates its camera to capture the user's hand movements in real time. The video data acquired by the camera is used as input and passed to the next process.

[0201] Step 2:

[0202] The terminal preprocesses the video data acquired from the camera using image processing equipment. Specifically, it uses the OpenCV library to convert the video data to grayscale, and then performs noise reduction and contrast adjustment. The processed image data is then output.

[0203] Step 3:

[0204] Based on the pre-processed image data, the terminal uses a feature point extraction mechanism to identify the positions of the joints and fingertips of the fingers. The identified feature point data is output and sent to the server.

[0205] Step 4:

[0206] The server receives feature point data transmitted from the terminal. Using the received feature point data as input, it analyzes the data using a generative AI model. As a result of the analysis, the gesture is identified.

[0207] Step 5:

[0208] The server generates corresponding commands based on the gesture recognition results analyzed by the generation AI model. In this case, commands such as "start," "stop," and "adjust" are generated depending on the recognized gesture. The generated commands are output and sent to the terminal.

[0209] Step 6:

[0210] The terminal receives commands sent from the server. It uses the received commands as input to execute specific robot operations. For example, upon receiving the "start" command, the terminal starts the conveyor belt. The result of this operation is output.

[0211] Step 7:

[0212] The terminal provides feedback to the user regarding the results of the operations performed. Specifically, it displays messages such as "Robot operation will begin" to the user. This allows the user to confirm that the operation was performed successfully.

[0213] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0214] This invention is a system that combines a system for recognizing the user's hand movements in real time and performing operations based on those movements with an emotion engine for recognizing the user's emotions and reflecting them in the feedback. This system consists of a camera, image processing, feature point extraction, server, gesture recognition, command interpretation and generation, terminal, and emotion engine.

[0215] System Configuration

[0216] Camera means

[0217] The camera built into the device captures the user's hand movements and facial expressions in real time. The captured image data is processed in real time.

[0218] Image processing means

[0219] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0220] Feature point extraction means

[0221] Feature points, such as the joints and fingertips of the fingers, are extracted from the pre-processed image data. This converts the shape and position of the fingers into numerical data.

[0222] Server Means

[0223] After receiving feature point data sent from the terminal, the server performs analysis using a machine learning model. This machine learning model has been pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[0224] Gesture recognition means

[0225] The server classifies the type of gesture based on the feature point data it analyzes. Based on this classification, it interprets and generates the relevant commands.

[0226] Command interpretation and generation means

[0227] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[0228] Emotional Engine

[0229] The server incorporates an emotion engine to analyze the user's facial expressions, voice, or other biometric data. This emotion engine recognizes the user's emotions and reflects them in the interpretation of gesture commands and the feedback provided.

[0230] Terminal means

[0231] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it might respond with something like, "Playing a calming song."

[0232] Specific example

[0233] When playing music using gestures

[0234] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera, captures the user's hand movements in real time, and simultaneously acquires their facial expressions.

[0235] 2. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[0236] 3. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[0237] 4. Based on this result, the server generates a music playback command and sends it to the terminal, along with a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[0238] 5. The device displays this feedback message, launches the music player application, and begins playing music.

[0239] This system allows users to intuitively perform various operations using only hand gestures, providing high operability and a superior user experience. Furthermore, by recognizing user emotions and responding accordingly, it achieves a more human-like interaction.

[0240] The following describes the processing flow.

[0241] Step 1:

[0242] The device activates its built-in camera. The camera begins capturing the user's hand movements and facial expressions in real time.

[0243] Step 2:

[0244] The device acquires captured images frame by frame in real time. The user's hand movements and facial expressions are continuously collected.

[0245] Step 3:

[0246] The device performs preprocessing on the acquired image data. Specifically, it converts the image to grayscale, applies a noise reduction filter, and adjusts the contrast. This improves the accuracy of subsequent processing.

[0247] Step 4:

[0248] The device segments the finger area from the pre-processed image. This process uses color distribution and contour detection algorithms.

[0249] Step 5:

[0250] The device extracts feature points from segmented finger regions. Specifically, it identifies the positions of finger joints and fingertips and records them as numerical data.

[0251] Step 6:

[0252] The device extracts facial feature points to analyze the user's facial expression data. This feature point data is analyzed in real time.

[0253] Step 7:

[0254] The device sends extracted finger and facial feature point data to the server. This transmission is done via the internet or a local network.

[0255] Step 8:

[0256] The server receives feature point data and prepares it for analysis. The received data is then loaded into the system.

[0257] Step 9:

[0258] The server uses machine learning models to analyze the feature point data of the fingers. This involves algorithms such as neural networks and support vector machines.

[0259] Step 10:

[0260] The server then uses an emotion engine to recognize the user's emotional state from facial feature point data. For example, it can determine "joy" from a smile and "anger" from wrinkles between the eyebrows.

[0261] Step 11:

[0262] The server recognizes the type of gesture (e.g., the "play" gesture) based on the analysis results of the hand feature point data.

[0263] Step 12:

[0264] The server interprets and generates specific commands (e.g., play music) based on recognized gestures and the user's emotional state. It also generates appropriate feedback messages tailored to the emotional state.

[0265] Step 13:

[0266] The server sends the generated command and feedback message (e.g., "Starting music playback," "Playing a song to refresh your mood") to the terminal.

[0267] Step 14:

[0268] The device displays a feedback message to the user. For example, it might display "Starting music playback" on the screen.

[0269] Step 15:

[0270] The terminal executes a command. Specifically, it launches a music player application and starts playing music. The type of music played is also adjusted according to the user's emotional state.

[0271] (Example 2)

[0272] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0273] There is a need to develop a system that can accurately recognize a user's hand movements and emotions in real time and perform actions based on that recognition. However, conventional technologies have insufficient accuracy in image processing, feature point extraction, and gesture recognition, and there are no systems that provide feedback that takes the user's emotional state into account. As a result, there are challenges such as a reduced user experience and difficulty in intuitive operation.

[0274] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0275] In this invention, the server includes a shooting device for acquiring the user's finger movements in real time, an image processing device for preprocessing image data acquired from the shooting device, a device for extracting feature points of the fingers from the image data processed by the image processing device, an analysis device for analyzing the feature point data, a device for recognizing gestures based on the feature point data analyzed by the analysis device, a device for interpreting and generating specific instructions based on the gesture recognition device, a terminal device for executing the instructions and providing feedback to the user, an emotion engine for analyzing the user's facial expression data and recognizing their emotional state, and a device for adjusting the feedback content based on the emotional state recognized by the emotion engine. This enables real-time, highly accurate gesture recognition and appropriate feedback that takes the user's emotions into consideration.

[0276] A "user" refers to a person who operates a system.

[0277] "Hand and finger movements" refers to the position and actions of the user's hands and fingers.

[0278] "Real-time" refers to processing or responding instantly without delay or waiting time.

[0279] "Image acquisition device" refers to a device such as a camera used to acquire images.

[0280] "Image data" refers to the digital information of an image acquired by a camera.

[0281] "Preprocessing" refers to the process of modifying image data to make it easier to analyze and extract feature points.

[0282] An "image processing device" refers to a device used for pre-processing image data.

[0283] "Feature points" refer to important points necessary for analysis, such as specific positions or joint points on the fingers.

[0284] "Extraction" refers to extracting feature points from image data.

[0285] "Analysis device" refers to a device that performs analysis based on feature point data.

[0286] "Gesture" refers to a specific action formed by combining the movements of fingers.

[0287] "Recognition" refers to discriminating gestures and emotions based on the analysis results.

[0288] "Instruction" refers to an action or command corresponding to a gesture.

[0289] "Terminal device" refers to an electronic device that a user operates or receives feedback from.

[0290] "Facial expression data" refers to data obtained by digitizing the movements and states of a user's face.

[0291] "Emotion engine" refers to a system that analyzes facial expression data, voice, etc. to recognize a user's emotions.

[0292] "Feedback" refers to the responses and information provided by the system to the user.

[0293] This invention relates to a system that recognizes a user's finger movements and emotions in real time and executes operations based on them. This system includes a photographing device, an image processing device, a feature point extraction device, an analysis device, a gesture recognition device, an instruction generation device, a terminal device, and an emotion engine.

[0294] Configuration of the system

[0295] Photographing device

[0296] When the user operates the system, the camera mounted on the terminal captures the movements of the fingers and expressions in real time. The captured image data is directly passed to the image processing device.

[0297] Image processing device

[0298] The terminal performs preprocessing on the image data acquired from the camera, such as grayscale conversion, noise removal, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0299] Feature point extraction device

[0300] Feature points such as the joints and tips of the fingers are extracted from the preprocessed image data. As a result, the shape and position of the fingers are converted into numerical data and transmitted to the analysis device.

[0301] Analysis device

[0302] The server that receives the feature point data transmitted from the terminal performs analysis using a machine learning model. This machine learning model has been trained in advance with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[0303] Gesture recognition device

[0304] Based on the feature point data analyzed by the analysis device, the type of gesture is classified. For example, it is recognized as a "play" gesture.

[0305] Instruction generation device

[0306] The analysis device generates an instruction corresponding to the recognized gesture. This instruction covers a wide range, such as launching a specific application, executing a specific operation, or displaying specific data.

[0307] Emotion engine

[0308] The server incorporates an emotion engine to analyze biometric data such as the user's facial expressions and voice. This emotion engine recognizes the user's emotions and reflects them in the interpretation of gesture-based instructions and feedback.

[0309] Terminal device

[0310] The terminal device receives instructions sent from the server and actually executes those instructions. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it will respond with something like, "Playing a calming song."

[0311] Specific example

[0312] When playing music using gestures

[0313] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera.

[0314] 2. The device activates its camera, capturing the user's hand movements in real time, and simultaneously acquiring their facial expressions.

[0315] 3. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[0316] 4. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[0317] 5. Based on this result, the server generates a music playback command and sends it to the terminal, including a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[0318] 6. The device displays this feedback message, launches the music player application, and begins playing music.

[0319] Example of a prompt

[0320] This system recognizes the user's hand movements and emotions in real time and performs actions based solely on that recognition. It captures hand movements and facial expressions with a camera, and then performs image processing and feature point extraction. The obtained data is sent to a server, which analyzes it to recognize gestures and emotions. The server then generates commands based on the recognition results and executes them on the terminal.

[0321] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0322] Step 1:

[0323] The user makes a hand gesture. The user performs a specific gesture (for example, a "play" gesture where the thumb and index finger draw a circle) in front of the device's camera. The camera captures this action.

[0324] Input: Finger movements

[0325] Output: Raw image data

[0326] Specific operation: The camera captures the user's hand movements in real time and generates image data.

[0327] Step 2:

[0328] The device performs image processing. The device performs preprocessing on image data acquired from the camera, such as grayscale conversion, noise reduction, and contrast adjustment. This processing improves the quality of the image data, enabling smoother subsequent feature point extraction.

[0329] Input: Raw image data

[0330] Output: Preprocessed image data

[0331] Specific operation: Converts a color image to grayscale, removes noise, and adjusts contrast.

[0332] Step 3:

[0333] The device extracts feature points. It extracts feature points such as the joints and fingertips of the fingers from the pre-processed image data. This converts the shape and position of the fingers into numerical data.

[0334] Input: Preprocessed image data

[0335] Output: Feature point data (numerical data)

[0336] Specific operation: Calculates the position of the joints and fingertips of the fingers and generates their coordinate data.

[0337] Step 4:

[0338] The terminal sends feature point data to the server. The terminal sends the extracted feature point data to the server.

[0339] Input: Feature point data

[0340] Output: Data transfer to the server

[0341] Specific operation: Send feature point data in JSON format to a specific server address.

[0342] Step 5:

[0343] The server performs gesture and emotion analysis. The server inputs the received feature point data into a machine learning model to recognize the type of gesture. Simultaneously, the emotion engine analyzes the user's facial expression data to recognize the user's emotional state.

[0344] Input: Feature point data, facial expression data

[0345] Output: Gesture recognition results, emotion recognition results

[0346] Specific operation: Feature point data is input into a machine learning model and recognized as a "play" gesture. Simultaneously, the emotion engine determines "joy" from the facial expression.

[0347] Step 6:

[0348] The server generates commands. The server generates appropriate commands based on the analysis of gestures and emotions. For example, it generates a music playback command for a "play" gesture.

[0349] Input: Gesture recognition results, emotion recognition results

[0350] Output: Music playback command

[0351] Specific operation: Combine the gesture recognition results and emotion recognition results to generate a music playback command in JSON format and send it to the terminal.

[0352] Step 7:

[0353] The server generates feedback messages. The server generates feedback messages that are tailored to the user's emotional state. For example, if the user is happy, it might generate a message such as, "We'll play some music to refresh your mood."

[0354] Input: Emotion recognition result

[0355] Output: Feedback message

[0356] Specific operation: Based on the user's emotional state, generate an appropriate feedback message and send it to the device along with instructions.

[0357] Step 8:

[0358] The terminal executes a command and displays feedback. The terminal executes a command sent from the server, for example, launching a music player application and starting music playback. Simultaneously, a feedback message is displayed to the user.

[0359] Input: Music playback command, feedback message

[0360] Output: Music playback, message display

[0361] Specific actions: Launch the music player and start playing music. Simultaneously, display the message "Playing music to refresh your mood" on the screen.

[0362] (Application Example 2)

[0363] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0364] In modern interactive systems, improving user operability and experience is a crucial challenge. However, conventional systems struggle to accurately recognize user hand movements and gestures, resulting in limitations in operation and a lack of features to provide feedback that responds to user emotions. This limits user interaction, and there is a growing demand for more human-like and flexible responses.

[0365] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0366] In this invention, the server includes: video acquisition means for acquiring the movements of the user's fingers in real time; image processing means for preprocessing image data acquired from the video acquisition means; means for extracting feature points of the fingers from the image data processed by the image processing means; analysis means for analyzing the feature point data; means for recognizing gestures based on the feature point data analyzed by the analysis means; means for interpreting and generating a specific command based on the gesture recognition means; execution means for executing the command and providing feedback of the result to the user; and emotion analysis means for analyzing the user's emotions when the command is executed and providing appropriate feedback. This makes it possible to accurately recognize the gestures of the user's fingers and provide appropriate feedback according to the user's emotions.

[0367] "Video acquisition means" refers to a device for acquiring the user's hand movements and facial expressions in real time.

[0368] "Image processing means" refers to means for pre-processing image data obtained from video acquisition means to improve its quality.

[0369] A "feature point extraction method" is a means of extracting feature points, such as the joints and fingertips of the fingers, from pre-processed image data.

[0370] "Analysis means" refers to a computer-based system that analyzes feature point data and recognizes gestures and emotions based on that analysis.

[0371] A "gesture recognition means" is a means of classifying a user's gestures based on feature point data recognized by an analysis means.

[0372] The "command interpretation and generation means" is a means for interpreting and generating commands corresponding to gestures recognized by the gesture recognition means.

[0373] "Execution means" refers to a means of executing commands generated by command interpretation and generation means and providing feedback of the execution results to the user.

[0374] "Emotion analysis methods" are means of recognizing a user's emotions by analyzing their facial expressions and voice data.

[0375] This invention is a system for recognizing the user's hand movements in real time and executing operations based on those movements, and further incorporates an emotion engine that recognizes the user's emotions and reflects them in the feedback. This system consists of video acquisition means, image processing means, feature point extraction means, analysis means, gesture recognition means, command interpretation and generation means, execution means, and emotion analysis means.

[0376] System Configuration

[0377] Video acquisition method

[0378] It includes a camera to capture the user's hand movements and facial expressions in real time. The camera is attached to a device (e.g., a smartphone or PC). This allows the user's hand movements and facial expressions to be collected.

[0379] Image processing means

[0380] It has a function to preprocess image data acquired from video acquisition devices. This preprocessing includes grayscale conversion, noise reduction, and contrast adjustment. The software used is an open-source image processing library (e.g., OpenCV). This preprocessing improves the quality of the image data and makes subsequent feature point extraction easier.

[0381] Feature point extraction means

[0382] Feature points, such as the joints and fingertips of the fingers, are extracted from pre-processed image data. This process converts the shape and position of the fingers into numerical data. The machine learning model used has been pre-trained on a large amount of gesture data.

[0383] Analysis means

[0384] The server receives feature point data sent from the terminal and analyzes this data using a machine learning model. This machine learning model uses a framework such as TensorFlow. The server analyzes the feature point data to recognize the type of gesture and the user's emotional state.

[0385] Gesture recognition means

[0386] The server identifies the type of gesture based on the analyzed feature point data. Based on this classification, it interprets and generates the relevant commands.

[0387] Command interpretation and generation means

[0388] The server generates commands corresponding to the recognized gestures. These commands are wide-ranging and include launching specific applications, performing specific actions, and displaying specific data.

[0389] Execution method

[0390] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it will respond with something like, "Playing a calming song."

[0391] Emotion analysis means

[0392] The device analyzes facial and voice data it collects to recognize the user's emotions. This emotion analysis utilizes libraries such as the EmotionRecognition library. The emotion analysis provides appropriate feedback for gesture interpretation and generated commands.

[0393] Specific example

[0394] When playing music using gestures

[0395] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera, captures the user's hand movements in real time, and simultaneously acquires their facial expressions.

[0396] 2. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[0397] 3. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[0398] 4. Based on this result, the server generates a music playback command and sends it to the terminal, along with a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[0399] 5. The device displays this feedback message, launches the music player application, and begins playing music.

[0400] Example of a prompt

[0401] "Implement an interactive system that recognizes the user's hand movements in real time and provides appropriate feedback based on the user's emotions. This system recognizes hand gestures using image processing and machine learning, and evaluates the user's emotions using an emotion engine."

[0402] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0403] Step 1:

[0404] The user performs a "play" gesture.

[0405] Input: User's hand and finger movements and facial expressions.

[0406] Action: The user makes a circular gesture with their thumb and index finger in front of the device's camera.

[0407] Step 2:

[0408] The device activates its camera and captures the user's hand movements and facial expressions in real time.

[0409] Input: Video data from the camera.

[0410] Output: Video data of hand and finger movements and facial expressions.

[0411] Operation: The camera records the user's gestures and facial expressions in real time.

[0412] Step 3:

[0413] The device performs pre-processing on the captured video data, including grayscale conversion, noise reduction, and contrast adjustment.

[0414] Input: Video data of hand and finger movements and facial expressions.

[0415] Output: Preprocessed image data.

[0416] Operation: Preprocessing is performed on the terminal side using image processing software (e.g., OpenCV) to improve the quality of the video data.

[0417] Step 4:

[0418] Feature points, such as the positions of finger joints and fingertips, are extracted from the pre-processed image data.

[0419] Input: Preprocessed image data.

[0420] Output: Feature point data for fingers.

[0421] Operation: The terminal uses a feature point extraction algorithm to convert the positions of finger joints and fingertips into numerical data.

[0422] Step 5:

[0423] The device sends characteristic point data of the fingers to the server.

[0424] Input: Feature point data for fingers.

[0425] Output: Sending data to the server.

[0426] Operation: The terminal sends the extracted feature point data to the server via the network.

[0427] Step 6:

[0428] The server analyzes the received feature point data using a machine learning model to recognize the type of gesture.

[0429] Input: Feature point data for fingers.

[0430] Output: Gesture recognition result.

[0431] Operation: Feature point data is analyzed using a machine learning framework (e.g., TensorFlow) running on the server side.

[0432] Step 7:

[0433] The server simultaneously analyzes facial expression data using an emotion engine to recognize the user's emotions.

[0434] Input: Facial expression data.

[0435] Output: Emotion recognition result.

[0436] Operation: The system uses an emotion analysis library (e.g., EmotionRecognition) running on the server side to analyze emotions from facial expression data.

[0437] Step 8:

[0438] The server generates a music playback command based on the gesture recognition results and emotion recognition results.

[0439] Input: Gesture recognition results, emotion recognition results.

[0440] Output: Music playback command.

[0441] Operation: The server interprets the user's intent based on the analysis results and generates an appropriate music playback command.

[0442] Step 9:

[0443] The terminal executes the command received from the server and starts playing music.

[0444] Input: Music playback command.

[0445] Output: Music playback begins.

[0446] Operation: The device launches the music player application and begins playing music.

[0447] Step 10:

[0448] The device displays feedback messages tailored to the user's emotional state.

[0449] Input: Emotion recognition result.

[0450] Output: Feedback message.

[0451] Operation: Based on the emotion recognition results, the device displays a feedback message on the screen that is appropriate for the user (e.g., "Playing a song to refresh your mood").

[0452] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0453] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0454] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0455] [Second Embodiment]

[0456] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0457] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0458] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0459] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0460] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0461] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0462] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0463] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0464] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0465] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0466] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0467] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0468] This invention is a system for recognizing the movements of a user's fingers in real time and performing operations based on those movements. This system consists of the following components: a camera, an image processing unit, a feature point extraction unit, a server, a gesture recognition unit, a command interpretation and generation unit, and a terminal unit.

[0469] System Configuration

[0470] Camera means

[0471] The camera built into the device captures the user's hand movements in real time. The captured image data is processed in real time.

[0472] Image processing means

[0473] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0474] Feature point extraction means

[0475] Feature points, such as the joints and fingertips of the fingers, are extracted from the preprocessed image data. This feature point extraction uses a specific algorithm to convert the shape and position of the fingers into numerical data.

[0476] Server Means

[0477] After receiving feature point data transmitted from the terminal, the server performs analysis using a machine learning model. This machine learning model has been pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[0478] Gesture recognition means

[0479] The server classifies the type of gesture based on the feature point data it analyzes. Based on this classification, it interprets and generates the relevant commands.

[0480] Command interpretation and generation means

[0481] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[0482] Terminal means

[0483] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback."

[0484] Specific example

[0485] When playing music using gestures

[0486] The user makes a "play" gesture (for example, drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera and captures the user's hand movements in real time. The acquired image is pre-processed by the device, including grayscale conversion, noise reduction, and contrast adjustment. Feature points of the hands are extracted from the pre-processed image, and this data is sent to the server.

[0487] The server uses a machine learning model to analyze feature point data and recognizes it as a "play" gesture. Based on this, the server generates a music playback command and sends it to the terminal. The terminal receives this command, launches the music player application, starts playing music, and provides feedback to the user saying, "Music playback will begin."

[0488] This system allows users to intuitively perform various operations using only hand gestures, providing high operability and a superior user experience.

[0489] The following describes the processing flow.

[0490] Step 1:

[0491] The device activates its built-in camera. The camera begins capturing the user's hand movements in real time.

[0492] Step 2:

[0493] The device acquires captured images frame by frame in real time. This allows for the continuous collection of the user's hand movements.

[0494] Step 3:

[0495] The device performs preprocessing on the acquired image data. Specifically, it converts the image to grayscale, applies a noise reduction filter, and adjusts the contrast. This improves the accuracy of subsequent processing.

[0496] Step 4:

[0497] The device segments the finger area from the pre-processed image. This is done using color distribution and contour detection algorithms.

[0498] Step 5:

[0499] The device extracts feature points from segmented finger regions. Specifically, it identifies the positions of finger joints and fingertips and records them as numerical data.

[0500] Step 6:

[0501] The terminal sends the extracted feature point data to the server. This transmission is done via the internet or a local network.

[0502] Step 7:

[0503] The server receives feature point data and prepares it for analysis. The received data is then loaded into the system.

[0504] Step 8:

[0505] The server analyzes the data using machine learning models. This involves algorithms such as neural networks and support vector machines.

[0506] Step 9:

[0507] The server recognizes specific gestures based on data analysis. For example, drawing a circle with the thumb and index finger is recognized as the "play" gesture.

[0508] Step 10:

[0509] The server interprets and generates a specific command (e.g., play music) based on the recognized gesture.

[0510] Step 11:

[0511] The server sends a feedback message (e.g., "Starting music playback") to the terminal along with the generated command.

[0512] Step 12:

[0513] The device displays a feedback message to the user. For example, it might display "Starting music playback" on the screen.

[0514] Step 13:

[0515] The terminal executes a command. Specifically, it launches the music player application and starts playing music.

[0516] (Example 1)

[0517] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0518] Currently, many users utilize smart devices to access a variety of applications and services. However, these operations are typically performed using touchscreens or physical buttons, often resulting in cumbersome user experiences. Furthermore, traditional methods make intuitive operation and the use of hand gestures difficult, making them particularly inconvenient when hands are dirty or when performing other tasks. A system is needed to solve these problems and provide a more intuitive and convenient way to operate the device.

[0519] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0520] In this invention, the server includes: a shooting device means for acquiring the user's finger movements in real time; an image conversion means for preprocessing image data acquired from the shooting device means; a means for extracting finger feature points from the image data processed by the image conversion means; an analysis device means for analyzing the feature point data; a means for recognizing gestures based on the feature point data analyzed by the analysis device means; a means for interpreting and generating specific commands based on the gesture recognition means; and a terminal device means for executing the commands and feeding the results back to the user. This allows the user to operate intuitively using finger movements, enabling touchless and convenient operation even while working.

[0521] "Photography device means" refers to a device for acquiring the movements of the user's fingers in real time, and includes cameras and the like.

[0522] The "image conversion means" is a means for preprocessing image data acquired from the imaging device means, and performs processing such as grayscale conversion, noise reduction, and contrast adjustment.

[0523] "Methods for extracting feature points of fingers" refer to methods for extracting feature points such as the joints and fingertips of fingers from preprocessed image data, and primarily utilize algorithms and deep learning models.

[0524] "Analysis device means" refers to a means of receiving feature point data transmitted from a terminal and performing analysis using a machine learning model.

[0525] "Means for recognizing gestures" refers to means for classifying and recognizing the type of gesture based on feature point data analyzed by an analysis device.

[0526] "Means for interpreting and generating commands" refers to a method for generating specific commands based on recognized gestures and for interpreting those commands.

[0527] A "terminal device means" is a means of receiving commands sent from a server, actually executing those commands, and providing feedback of the results to the user.

[0528] This invention is a system for recognizing the movements of a user's fingers in real time and performing operations based on those movements. This system consists of the following components: a shooting device, an image conversion device, a device for extracting characteristic points of the fingers, an analysis device, a gesture recognition device, a command interpretation and generation device, and a terminal device.

[0529] System Configuration

[0530] Photography device means

[0531] The camera built into the device captures the user's hand movements in real time. The captured image data is processed in real time.

[0532] Image conversion means

[0533] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. For example, the OpenCV library is used for this preprocessing. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0534] A method for extracting characteristic points of the fingers.

[0535] Feature points, such as the joints and fingertips of the fingers, are extracted from the preprocessed image data. For this feature point extraction, a model like the MediaPipe Hand is used. The feature points are obtained as numerical data representing the shape and position of the fingers.

[0536] analysis equipment means

[0537] After receiving feature point data sent from the terminal, the server performs analysis using a machine learning model. This machine learning model is, for example, a gesture recognition model trained with TensorFlow. The server uses this model to recognize the type of gesture based on the newly received data.

[0538] Means of recognizing gestures

[0539] The server classifies the gesture type based on the analysis results. Based on this classification, it interprets and generates the relevant commands.

[0540] Means for interpreting and generating commands

[0541] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[0542] Terminal device means

[0543] The device receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, the device launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback."

[0544] Specific example

[0545] When playing music using gestures

[0546] The user makes a "play" gesture (for example, drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera and captures the user's hand movements in real time. The device preprocesses the acquired image by performing grayscale conversion, noise reduction, contrast adjustment, etc. Feature points of the hands are extracted from the preprocessed image and this data is sent to the server. The server analyzes the feature point data using a machine learning model and recognizes that it is a "play" gesture. Based on this result, the server generates a music playback command and sends it to the device. The device receives this command, launches the music player application, starts music playback, and provides feedback to the user saying, "Music playback will begin."

[0547] This system allows users to intuitively perform various operations using only hand gestures. An example of a prompt to be input to the AI ​​model is as follows:

[0548] Example of a prompt

[0549] Prompt: Create a system that recognizes a user's finger gesture for "play" and plays music accordingly. The steps are as follows:

[0550] 1. Use the device's camera to record the movement of your fingers.

[0551] 2. Preprocess the acquired image data (grayscale conversion, noise reduction, contrast adjustment).

[0552] 3. Extract feature points from the preprocessed image.

[0553] 4. Send the extracted feature point data to the server.

[0554] 5. The server analyzes the feature point data and recognizes the gesture.

[0555] 6. Generate a command based on the recognition result and send it to the terminal.

[0556] 7. The terminal receives the command and plays the music.

[0557] This allows users to operate the system with high operability and an intuitive interface.

[0558] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0559] Step 1:

[0560] Camera activation and image capture

[0561] The device activates its built-in camera and captures the user's finger movements in real time. Suppose the user performs a "play" gesture in front of the camera (for example, drawing a circle with their thumb and index finger). The input is the user's finger movements, and the output is real-time image data. This image data is retained for subsequent processing.

[0562] Specific actions:

[0563] The user moves their fingers in front of the device, and the device activates its camera to take a picture.

[0564] Step 2:

[0565] Image preprocessing

[0566] The device performs preprocessing on image data acquired in real time, including grayscale conversion, noise reduction, and contrast adjustment. This improves the quality of the image data and facilitates subsequent processing. The input is the image data from step 1, and the output is the preprocessed image data.

[0567] Specific actions:

[0568] The device converts the image data to grayscale, applies a noise reduction filter, and adjusts the image contrast.

[0569] Step 3:

[0570] Feature point extraction

[0571] The device extracts feature points of the fingers from pre-processed image data. For feature point extraction, for example, the MediaPipe Hand model is used. This model converts the positions of finger joints and fingertips into numerical data. Pre-processed image data is taken as input, and feature point data of the fingers is generated as output.

[0572] Specific actions:

[0573] The device uses the MediaPipe Hand model to calculate the positions of the joints and fingertips of the fingers and obtains numerical data indicating these positions.

[0574] Step 4:

[0575] Sending feature point data to the server

[0576] The device sends feature point data of the fingers to the server. This data is sent to the server via the network as a POST request. The input is feature point data, and the output is a request sent to the server.

[0577] Specific actions:

[0578] The terminal converts the feature point data into JSON format and sends a POST request to the server's API endpoint.

[0579] Step 5:

[0580] Gesture recognition and classification

[0581] The server uses the received feature point data to analyze it with a machine learning model and recognize the type of gesture. The model used is, for example, a gesture recognition model trained with TensorFlow. The input is feature point data, and the output is the gesture recognition result.

[0582] Specific actions:

[0583] The server inputs the received feature point data into a TensorFlow model and performs the analysis. The model recognizes the "play" gesture and generates the recognition result.

[0584] Step 6:

[0585] Command generation and sending

[0586] The server generates a command corresponding to the recognized gesture and sends that command to the terminal. This command could include, for example, an instruction to play music. The input is the gesture recognition result, and the output is the generated command.

[0587] Specific actions:

[0588] The server analyzes the gesture recognition results and generates a corresponding music playback command. The generated command is then sent to the terminal in JSON format.

[0589] Step 7:

[0590] Command execution and feedback

[0591] The terminal receives and executes a command sent from the server. It launches a music player application, starts playing music, and provides feedback to the user saying, "Starting music playback." The input is a command, and the output is music playback and a feedback message.

[0592] Specific actions:

[0593] The terminal analyzes the music playback command, launches the music player, and starts playback. At the same time, a message saying "Music playback will begin" is displayed on the screen to the user.

[0594] (Application Example 1)

[0595] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0596] In modern factories, improving work efficiency and ensuring worker safety are critical challenges. However, conventional robot operating systems often lack intuitive and rapid response capabilities, potentially leading to operational errors and increased worker burden. In particular, the presence of numerous buttons and switches increases the complexity of manual operation and makes operational errors more likely. To address these challenges, a system is needed that more intuitively recognizes worker movements and provides accurate operational feedback.

[0597] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0598] In this invention, the server includes a camera means for acquiring the user's finger movements in real time, an image processing means for preprocessing image data acquired from the camera means, a means for extracting finger feature points from the image data processed by the image processing means, a server means for analyzing the feature point data, a means for recognizing gestures based on the feature point data analyzed by the server means, a means for interpreting and generating specific commands based on the gesture recognition means, a terminal means for executing the commands and feeding the results back to the user, a means for operating the robot in a specific work environment, and a means for starting, stopping, or adjusting work based on gestures using the means. This enables the operator to intuitively and quickly operate the robot using only finger movements, thereby improving work efficiency and ensuring safety.

[0599] A "camera device" is a device used to capture the movements of the user's fingers in real time, and its role is to acquire image data.

[0600] "Image processing means" refers to techniques used to preprocess image data acquired from a camera and improve the quality of the processed image data.

[0601] The "feature point extraction means" is an algorithm for extracting feature points of fingers from pre-processed image data and converting their position and shape into numerical data.

[0602] A "server system" is a computer server that receives feature point data transmitted from a terminal and analyzes the data using a machine learning model.

[0603] "Gesture recognition means" refers to a technology that classifies and recognizes the type of user gesture based on feature point data analyzed by the server means.

[0604] The "command interpretation and generation means" is a means for interpreting and generating a corresponding command in response to a gesture recognized by the gesture recognition means.

[0605] "Terminal means" refers to equipment or devices that execute commands sent from server means and provide feedback of the results to the user.

[0606] "Means for operating robots in a specific work environment" refers to means used to direct specific actions or tasks when operating robots in a factory or work environment.

[0607] "Means for starting, stopping, or adjusting work based on gestures" refers to technology that recognizes the user's hand gestures and starts, stops, or adjusts the robot's movements in response to those gestures.

[0608] This invention is a system for intuitively operating a robot using specific gestures in a factory work environment. Detailed embodiments of this system are described below.

[0609] Key components of the system

[0610] Camera means

[0611] A camera system is used to capture the user's hand movements in real time. The camera, using a webcam or a dedicated high-resolution camera, captures the hand movements and acquires the data.

[0612] Image processing means

[0613] The acquired image data is preprocessed by image processing equipment. Specifically, grayscale conversion, noise reduction, and contrast adjustment are performed using image processing libraries such as OpenCV. This improves the quality of the image data and facilitates subsequent processing by feature point extraction equipment.

[0614] Feature point extraction means

[0615] Feature points, such as the joints and fingertip positions of the fingers, are extracted from the preprocessed image data. Specific algorithms (e.g., HOG or SIFT) are used to convert the shape and position of the fingers into numerical data.

[0616] Server Means

[0617] Feature point data is sent from the terminal to the server. The server analyzes the feature point data using a machine learning model. This machine learning model is built using frameworks such as TensorFlow and trained on a large amount of gesture data.

[0618] Gesture recognition means

[0619] The server classifies the type of gesture based on the analyzed feature point data. For example, if a user makes a "start" gesture, the server will recognize that gesture correctly.

[0620] Command interpretation and generation means

[0621] Based on the gesture recognition results, the server generates corresponding commands. Specifically, commands such as "start," "stop," and "adjust" are generated.

[0622] Terminal means

[0623] The generated command is sent to the terminal, which then executes the command. For example, if the command "start" is received, the robot will begin operation. The terminal also provides feedback to the user, such as "The robot will begin its work."

[0624] System operation example

[0625] For example, recognizing the "start" gesture will cause the conveyor belt in the factory to begin moving. Recognizing the "stop" gesture will cause the conveyor belt to stop. Recognizing the "adjust" gesture will adjust the robot's speed and direction. In this way, users can intuitively and quickly operate robots using only finger movements.

[0626] Example of a prompt

[0627] The following are examples of prompts to input into a generative AI model in a real-world usage scenario:

[0628] When the "start" gesture is recognized, the conveyor belts inside the factory begin to move.

[0629] When the "stop" gesture is recognized, the conveyor belt will stop.

[0630] When it recognizes the "adjust" gesture, it adjusts the robot's movement speed and direction.

[0631] As described above, the present invention improves work efficiency and worker safety within factories.

[0632] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0633] Step 1:

[0634] The terminal activates its camera to capture the user's hand movements in real time. The video data acquired by the camera is used as input and passed to the next process.

[0635] Step 2:

[0636] The terminal preprocesses the video data acquired from the camera using image processing equipment. Specifically, it uses the OpenCV library to convert the video data to grayscale, and then performs noise reduction and contrast adjustment. The processed image data is then output.

[0637] Step 3:

[0638] Based on the pre-processed image data, the terminal uses a feature point extraction mechanism to identify the positions of the joints and fingertips of the fingers. The identified feature point data is output and sent to the server.

[0639] Step 4:

[0640] The server receives feature point data transmitted from the terminal. Using the received feature point data as input, it analyzes the data using a generative AI model. As a result of the analysis, the gesture is identified.

[0641] Step 5:

[0642] The server generates corresponding commands based on the gesture recognition results analyzed by the generation AI model. In this case, commands such as "start," "stop," and "adjust" are generated depending on the recognized gesture. The generated commands are output and sent to the terminal.

[0643] Step 6:

[0644] The terminal receives commands sent from the server. It uses the received commands as input to execute specific robot operations. For example, upon receiving the "start" command, the terminal starts the conveyor belt. The result of this operation is output.

[0645] Step 7:

[0646] The terminal provides feedback to the user regarding the results of the operations performed. Specifically, it displays messages such as "Robot operation will begin" to the user. This allows the user to confirm that the operation was performed successfully.

[0647] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0648] This invention is a system that combines a system for recognizing the user's hand movements in real time and performing operations based on those movements with an emotion engine for recognizing the user's emotions and reflecting them in the feedback. This system consists of a camera, image processing, feature point extraction, server, gesture recognition, command interpretation and generation, terminal, and emotion engine.

[0649] System Configuration

[0650] Camera means

[0651] The camera built into the device captures the user's hand movements and facial expressions in real time. The captured image data is processed in real time.

[0652] Image processing means

[0653] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0654] Feature point extraction means

[0655] Feature points, such as the joints and fingertips of the fingers, are extracted from the pre-processed image data. This converts the shape and position of the fingers into numerical data.

[0656] Server Means

[0657] After receiving feature point data sent from the terminal, the server performs analysis using a machine learning model. This machine learning model has been pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[0658] Gesture recognition means

[0659] The server classifies the type of gesture based on the feature point data it analyzes. Based on this classification, it interprets and generates the relevant commands.

[0660] Command interpretation and generation means

[0661] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[0662] Emotional Engine

[0663] The server incorporates an emotion engine to analyze the user's facial expressions, voice, or other biometric data. This emotion engine recognizes the user's emotions and reflects them in the interpretation of gesture commands and the feedback provided.

[0664] Terminal means

[0665] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it might respond with something like, "Playing a calming song."

[0666] Specific example

[0667] When playing music using gestures

[0668] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera, captures the user's hand movements in real time, and simultaneously acquires their facial expressions.

[0669] 2. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[0670] 3. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[0671] 4. Based on this result, the server generates a music playback command and sends it to the terminal, along with a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[0672] 5. The device displays this feedback message, launches the music player application, and begins playing music.

[0673] This system allows users to intuitively perform various operations using only hand gestures, providing high operability and a superior user experience. Furthermore, by recognizing user emotions and responding accordingly, it achieves a more human-like interaction.

[0674] The following describes the processing flow.

[0675] Step 1:

[0676] The device activates its built-in camera. The camera begins capturing the user's hand movements and facial expressions in real time.

[0677] Step 2:

[0678] The device acquires captured images frame by frame in real time. The user's hand movements and facial expressions are continuously collected.

[0679] Step 3:

[0680] The device performs preprocessing on the acquired image data. Specifically, it converts the image to grayscale, applies a noise reduction filter, and adjusts the contrast. This improves the accuracy of subsequent processing.

[0681] Step 4:

[0682] The device segments the finger area from the pre-processed image. This process uses color distribution and contour detection algorithms.

[0683] Step 5:

[0684] The device extracts feature points from segmented finger regions. Specifically, it identifies the positions of finger joints and fingertips and records them as numerical data.

[0685] Step 6:

[0686] The device extracts facial feature points to analyze the user's facial expression data. This feature point data is analyzed in real time.

[0687] Step 7:

[0688] The device sends extracted finger and facial feature point data to the server. This transmission is done via the internet or a local network.

[0689] Step 8:

[0690] The server receives feature point data and prepares it for analysis. The received data is then loaded into the system.

[0691] Step 9:

[0692] The server uses machine learning models to analyze the feature point data of the fingers. This involves algorithms such as neural networks and support vector machines.

[0693] Step 10:

[0694] The server then uses an emotion engine to recognize the user's emotional state from facial feature point data. For example, it can determine "joy" from a smile and "anger" from wrinkles between the eyebrows.

[0695] Step 11:

[0696] The server recognizes the type of gesture (e.g., the "play" gesture) based on the analysis results of the hand feature point data.

[0697] Step 12:

[0698] The server interprets and generates specific commands (e.g., play music) based on recognized gestures and the user's emotional state. It also generates appropriate feedback messages tailored to the emotional state.

[0699] Step 13:

[0700] The server sends the generated command and feedback message (e.g., "Starting music playback," "Playing a song to refresh your mood") to the terminal.

[0701] Step 14:

[0702] The device displays a feedback message to the user. For example, it might display "Starting music playback" on the screen.

[0703] Step 15:

[0704] The terminal executes a command. Specifically, it launches a music player application and starts playing music. The type of music played is also adjusted according to the user's emotional state.

[0705] (Example 2)

[0706] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0707] There is a need to develop a system that can accurately recognize a user's hand movements and emotions in real time and perform actions based on that recognition. However, conventional technologies have insufficient accuracy in image processing, feature point extraction, and gesture recognition, and there are no systems that provide feedback that takes the user's emotional state into account. As a result, there are challenges such as a reduced user experience and difficulty in intuitive operation.

[0708] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0709] In this invention, the server includes a shooting device for acquiring the user's finger movements in real time, an image processing device for preprocessing image data acquired from the shooting device, a device for extracting feature points of the fingers from the image data processed by the image processing device, an analysis device for analyzing the feature point data, a device for recognizing gestures based on the feature point data analyzed by the analysis device, a device for interpreting and generating specific instructions based on the gesture recognition device, a terminal device for executing the instructions and providing feedback to the user, an emotion engine for analyzing the user's facial expression data and recognizing their emotional state, and a device for adjusting the feedback content based on the emotional state recognized by the emotion engine. This enables real-time, highly accurate gesture recognition and appropriate feedback that takes the user's emotions into consideration.

[0710] A "user" refers to a person who operates a system.

[0711] "Hand and finger movements" refers to the position and actions of the user's hands and fingers.

[0712] "Real-time" refers to processing or responding instantly without delay or waiting time.

[0713] "Image acquisition device" refers to a device such as a camera used to acquire images.

[0714] "Image data" refers to the digital information of an image acquired by a camera.

[0715] "Preprocessing" refers to the process of modifying image data to make it easier to analyze and extract feature points.

[0716] An "image processing device" refers to a device used for pre-processing image data.

[0717] "Feature points" refer to important points necessary for analysis, such as specific positions or joint points on the fingers.

[0718] "Extraction" refers to the process of extracting feature points from image data.

[0719] An "analysis device" refers to a device that performs analysis based on feature point data.

[0720] A "gesture" refers to a specific action formed by combining hand and finger movements.

[0721] "Recognition" refers to the process of identifying gestures and emotions based on the analysis results.

[0722] "Instructions" refer to actions or commands that correspond to gestures.

[0723] A "terminal device" refers to an electronic device that a user operates or uses to receive feedback.

[0724] "Facial expression data" refers to digital information obtained from the movements and state of a user's face.

[0725] An "emotion engine" refers to a system that analyzes facial expression data, voice, and other information to recognize a user's emotions.

[0726] "Feedback" refers to the responses or information that a system provides to a user.

[0727] This invention relates to a system that recognizes a user's hand movements and emotions in real time and performs operations based on them. The system includes a camera, an image processing device, a feature point extraction device, an analysis device, a gesture recognition device, an instruction generation device, a terminal device, and an emotion engine.

[0728] System Configuration

[0729] Imaging device

[0730] When a user operates the system, a camera on the terminal captures their hand movements and facial expressions in real time. The captured image data is then directly transmitted to an image processing device.

[0731] Image processing device

[0732] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0733] Feature point extraction device

[0734] Feature points, such as the joints and fingertips of the fingers, are extracted from the pre-processed image data. This converts the shape and position of the fingers into numerical data, which is then transmitted to the analysis device.

[0735] analysis device

[0736] A server receives feature point data sent from a terminal and performs analysis using a machine learning model. This machine learning model is pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[0737] Gesture recognition device

[0738] The analysis device classifies the type of gesture based on the feature point data it analyzes. For example, it might be recognized as a "play" gesture.

[0739] instruction generation device

[0740] The analysis device generates instructions corresponding to the recognized gesture. These instructions can range from launching a specific application, performing a specific action, to displaying specific data.

[0741] Emotional Engine

[0742] The server incorporates an emotion engine to analyze biometric data such as the user's facial expressions and voice. This emotion engine recognizes the user's emotions and reflects them in the interpretation of gesture-based instructions and feedback.

[0743] Terminal device

[0744] The terminal device receives instructions sent from the server and actually executes those instructions. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it will respond with something like, "Playing a calming song."

[0745] Specific example

[0746] When playing music using gestures

[0747] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera.

[0748] 2. The device activates its camera, capturing the user's hand movements in real time, and simultaneously acquiring their facial expressions.

[0749] 3. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[0750] 4. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[0751] 5. Based on this result, the server generates a music playback command and sends it to the terminal, including a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[0752] 6. The device displays this feedback message, launches the music player application, and begins playing music.

[0753] Example of a prompt

[0754] This system recognizes the user's hand movements and emotions in real time and performs actions based solely on that recognition. It captures hand movements and facial expressions with a camera, and then performs image processing and feature point extraction. The obtained data is sent to a server, which analyzes it to recognize gestures and emotions. The server then generates commands based on the recognition results and executes them on the terminal.

[0755] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0756] Step 1:

[0757] The user makes a hand gesture. The user performs a specific gesture (for example, a "play" gesture where the thumb and index finger draw a circle) in front of the device's camera. The camera captures this action.

[0758] Input: Finger movements

[0759] Output: Raw image data

[0760] Specific operation: The camera captures the user's hand movements in real time and generates image data.

[0761] Step 2:

[0762] The device performs image processing. The device performs preprocessing on image data acquired from the camera, such as grayscale conversion, noise reduction, and contrast adjustment. This processing improves the quality of the image data, enabling smoother subsequent feature point extraction.

[0763] Input: Raw image data

[0764] Output: Preprocessed image data

[0765] Specific operation: Converts a color image to grayscale, removes noise, and adjusts contrast.

[0766] Step 3:

[0767] The device extracts feature points. It extracts feature points such as the joints and fingertips of the fingers from the pre-processed image data. This converts the shape and position of the fingers into numerical data.

[0768] Input: Preprocessed image data

[0769] Output: Feature point data (numerical data)

[0770] Specific operation: Calculates the position of the joints and fingertips of the fingers and generates their coordinate data.

[0771] Step 4:

[0772] The terminal sends feature point data to the server. The terminal sends the extracted feature point data to the server.

[0773] Input: Feature point data

[0774] Output: Data transfer to the server

[0775] Specific operation: Send feature point data in JSON format to a specific server address.

[0776] Step 5:

[0777] The server performs gesture and emotion analysis. The server inputs the received feature point data into a machine learning model to recognize the type of gesture. Simultaneously, the emotion engine analyzes the user's facial expression data to recognize the user's emotional state.

[0778] Input: Feature point data, facial expression data

[0779] Output: Gesture recognition results, emotion recognition results

[0780] Specific operation: Feature point data is input into a machine learning model and recognized as a "play" gesture. Simultaneously, the emotion engine determines "joy" from the facial expression.

[0781] Step 6:

[0782] The server generates commands. The server generates appropriate commands based on the analysis of gestures and emotions. For example, it generates a music playback command for a "play" gesture.

[0783] Input: Gesture recognition results, emotion recognition results

[0784] Output: Music playback command

[0785] Specific operation: Combine the gesture recognition results and emotion recognition results to generate a music playback command in JSON format and send it to the terminal.

[0786] Step 7:

[0787] The server generates feedback messages. The server generates feedback messages that are tailored to the user's emotional state. For example, if the user is happy, it might generate a message such as, "We'll play some music to refresh your mood."

[0788] Input: Emotion recognition result

[0789] Output: Feedback message

[0790] Specific operation: Based on the user's emotional state, generate an appropriate feedback message and send it to the device along with instructions.

[0791] Step 8:

[0792] The terminal executes a command and displays feedback. The terminal executes a command sent from the server, for example, launching a music player application and starting music playback. Simultaneously, a feedback message is displayed to the user.

[0793] Input: Music playback command, feedback message

[0794] Output: Music playback, message display

[0795] Specific actions: Launch the music player and start playing music. Simultaneously, display the message "Playing music to refresh your mood" on the screen.

[0796] (Application Example 2)

[0797] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0798] In modern interactive systems, improving user operability and experience is a crucial challenge. However, conventional systems struggle to accurately recognize user hand movements and gestures, resulting in limitations in operation and a lack of features to provide feedback that responds to user emotions. This limits user interaction, and there is a growing demand for more human-like and flexible responses.

[0799] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0800] In this invention, the server includes: video acquisition means for acquiring the movements of the user's fingers in real time; image processing means for preprocessing image data acquired from the video acquisition means; means for extracting feature points of the fingers from the image data processed by the image processing means; analysis means for analyzing the feature point data; means for recognizing gestures based on the feature point data analyzed by the analysis means; means for interpreting and generating a specific command based on the gesture recognition means; execution means for executing the command and providing feedback of the result to the user; and emotion analysis means for analyzing the user's emotions when the command is executed and providing appropriate feedback. This makes it possible to accurately recognize the gestures of the user's fingers and provide appropriate feedback according to the user's emotions.

[0801] "Video acquisition means" refers to a device for acquiring the user's hand movements and facial expressions in real time.

[0802] "Image processing means" refers to means for pre-processing image data obtained from video acquisition means to improve its quality.

[0803] A "feature point extraction method" is a means of extracting feature points, such as the joints and fingertips of the fingers, from pre-processed image data.

[0804] "Analysis means" refers to a computer-based system that analyzes feature point data and recognizes gestures and emotions based on that analysis.

[0805] A "gesture recognition means" is a means of classifying a user's gestures based on feature point data recognized by an analysis means.

[0806] The "command interpretation and generation means" is a means for interpreting and generating commands corresponding to gestures recognized by the gesture recognition means.

[0807] "Execution means" refers to a means of executing commands generated by command interpretation and generation means and providing feedback of the execution results to the user.

[0808] "Emotion analysis methods" are means of recognizing a user's emotions by analyzing their facial expressions and voice data.

[0809] This invention is a system for recognizing the user's hand movements in real time and executing operations based on those movements, and further incorporates an emotion engine that recognizes the user's emotions and reflects them in the feedback. This system consists of video acquisition means, image processing means, feature point extraction means, analysis means, gesture recognition means, command interpretation and generation means, execution means, and emotion analysis means.

[0810] System Configuration

[0811] Video acquisition method

[0812] It includes a camera to capture the user's hand movements and facial expressions in real time. The camera is attached to a device (e.g., a smartphone or PC). This allows the user's hand movements and facial expressions to be collected.

[0813] Image processing means

[0814] It has a function to preprocess image data acquired from video acquisition devices. This preprocessing includes grayscale conversion, noise reduction, and contrast adjustment. The software used is an open-source image processing library (e.g., OpenCV). This preprocessing improves the quality of the image data and makes subsequent feature point extraction easier.

[0815] Feature point extraction means

[0816] Feature points, such as the joints and fingertips of the fingers, are extracted from pre-processed image data. This process converts the shape and position of the fingers into numerical data. The machine learning model used has been pre-trained on a large amount of gesture data.

[0817] Analysis means

[0818] The server receives feature point data sent from the terminal and analyzes this data using a machine learning model. This machine learning model uses a framework such as TensorFlow. The server analyzes the feature point data to recognize the type of gesture and the user's emotional state.

[0819] Gesture recognition means

[0820] The server identifies the type of gesture based on the analyzed feature point data. Based on this classification, it interprets and generates the relevant commands.

[0821] Command interpretation and generation means

[0822] The server generates commands corresponding to the recognized gestures. These commands are wide-ranging and include launching specific applications, performing specific actions, and displaying specific data.

[0823] Execution method

[0824] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it will respond with something like, "Playing a calming song."

[0825] Emotion analysis means

[0826] The device analyzes facial and voice data it collects to recognize the user's emotions. This emotion analysis utilizes libraries such as the EmotionRecognition library. The emotion analysis provides appropriate feedback for gesture interpretation and generated commands.

[0827] Specific example

[0828] When playing music using gestures

[0829] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera, captures the user's hand movements in real time, and simultaneously acquires their facial expressions.

[0830] 2. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[0831] 3. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[0832] 4. Based on this result, the server generates a music playback command and sends it to the terminal, along with a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[0833] 5. The device displays this feedback message, launches the music player application, and begins playing music.

[0834] Example of a prompt

[0835] "Implement an interactive system that recognizes the user's hand movements in real time and provides appropriate feedback based on the user's emotions. This system recognizes hand gestures using image processing and machine learning, and evaluates the user's emotions using an emotion engine."

[0836] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0837] Step 1:

[0838] The user performs a "play" gesture.

[0839] Input: User's hand and finger movements and facial expressions.

[0840] Action: The user makes a circular gesture with their thumb and index finger in front of the device's camera.

[0841] Step 2:

[0842] The device activates its camera and captures the user's hand movements and facial expressions in real time.

[0843] Input: Video data from the camera.

[0844] Output: Video data of hand and finger movements and facial expressions.

[0845] Operation: The camera records the user's gestures and facial expressions in real time.

[0846] Step 3:

[0847] The device performs pre-processing on the captured video data, including grayscale conversion, noise reduction, and contrast adjustment.

[0848] Input: Video data of hand and finger movements and facial expressions.

[0849] Output: Preprocessed image data.

[0850] Operation: Preprocessing is performed on the terminal side using image processing software (e.g., OpenCV) to improve the quality of the video data.

[0851] Step 4:

[0852] Feature points, such as the positions of finger joints and fingertips, are extracted from the pre-processed image data.

[0853] Input: Preprocessed image data.

[0854] Output: Feature point data for fingers.

[0855] Operation: The terminal uses a feature point extraction algorithm to convert the positions of finger joints and fingertips into numerical data.

[0856] Step 5:

[0857] The device sends characteristic point data of the fingers to the server.

[0858] Input: Feature point data for fingers.

[0859] Output: Sending data to the server.

[0860] Operation: The terminal sends the extracted feature point data to the server via the network.

[0861] Step 6:

[0862] The server analyzes the received feature point data using a machine learning model to recognize the type of gesture.

[0863] Input: Feature point data for fingers.

[0864] Output: Gesture recognition result.

[0865] Operation: Feature point data is analyzed using a machine learning framework (e.g., TensorFlow) running on the server side.

[0866] Step 7:

[0867] The server simultaneously analyzes facial expression data using an emotion engine to recognize the user's emotions.

[0868] Input: Facial expression data.

[0869] Output: Emotion recognition result.

[0870] Operation: The system uses an emotion analysis library (e.g., EmotionRecognition) running on the server side to analyze emotions from facial expression data.

[0871] Step 8:

[0872] The server generates a music playback command based on the gesture recognition results and emotion recognition results.

[0873] Input: Gesture recognition results, emotion recognition results.

[0874] Output: Music playback command.

[0875] Operation: The server interprets the user's intent based on the analysis results and generates an appropriate music playback command.

[0876] Step 9:

[0877] The terminal executes the command received from the server and starts playing music.

[0878] Input: Music playback command.

[0879] Output: Music playback begins.

[0880] Operation: The device launches the music player application and begins playing music.

[0881] Step 10:

[0882] The device displays feedback messages tailored to the user's emotional state.

[0883] Input: Emotion recognition result.

[0884] Output: Feedback message.

[0885] Operation: Based on the emotion recognition results, the device displays a feedback message on the screen that is appropriate for the user (e.g., "Playing a song to refresh your mood").

[0886] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0887] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0888] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0889] [Third Embodiment]

[0890] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0891] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0892] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0893] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0894] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0895] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0896] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0897] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0898] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0899] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0900] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0901] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0902] This invention is a system for recognizing the movements of a user's fingers in real time and performing operations based on those movements. This system consists of the following components: a camera, an image processing unit, a feature point extraction unit, a server, a gesture recognition unit, a command interpretation and generation unit, and a terminal unit.

[0903] System Configuration

[0904] Camera means

[0905] The camera built into the device captures the user's hand movements in real time. The captured image data is processed in real time.

[0906] Image processing means

[0907] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0908] Feature point extraction means

[0909] Feature points, such as the joints and fingertips of the fingers, are extracted from the preprocessed image data. This feature point extraction uses a specific algorithm to convert the shape and position of the fingers into numerical data.

[0910] Server Means

[0911] After receiving feature point data transmitted from the terminal, the server performs analysis using a machine learning model. This machine learning model has been pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[0912] Gesture recognition means

[0913] The server classifies the type of gesture based on the feature point data it analyzes. Based on this classification, it interprets and generates the relevant commands.

[0914] Command interpretation and generation means

[0915] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[0916] Terminal means

[0917] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback."

[0918] Specific example

[0919] When playing music using gestures

[0920] The user makes a "play" gesture (for example, drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera and captures the user's hand movements in real time. The acquired image is pre-processed by the device, including grayscale conversion, noise reduction, and contrast adjustment. Feature points of the hands are extracted from the pre-processed image, and this data is sent to the server.

[0921] The server uses a machine learning model to analyze feature point data and recognizes it as a "play" gesture. Based on this, the server generates a music playback command and sends it to the terminal. The terminal receives this command, launches the music player application, starts playing music, and provides feedback to the user saying, "Music playback will begin."

[0922] This system allows users to intuitively perform various operations using only hand gestures, providing high operability and a superior user experience.

[0923] The following describes the processing flow.

[0924] Step 1:

[0925] The device activates its built-in camera. The camera begins capturing the user's hand movements in real time.

[0926] Step 2:

[0927] The device acquires captured images frame by frame in real time. This allows for the continuous collection of the user's hand movements.

[0928] Step 3:

[0929] The device performs preprocessing on the acquired image data. Specifically, it converts the image to grayscale, applies a noise reduction filter, and adjusts the contrast. This improves the accuracy of subsequent processing.

[0930] Step 4:

[0931] The device segments the finger area from the pre-processed image. This is done using color distribution and contour detection algorithms.

[0932] Step 5:

[0933] The device extracts feature points from segmented finger regions. Specifically, it identifies the positions of finger joints and fingertips and records them as numerical data.

[0934] Step 6:

[0935] The terminal sends the extracted feature point data to the server. This transmission is done via the internet or a local network.

[0936] Step 7:

[0937] The server receives feature point data and prepares it for analysis. The received data is then loaded into the system.

[0938] Step 8:

[0939] The server analyzes the data using machine learning models. This involves algorithms such as neural networks and support vector machines.

[0940] Step 9:

[0941] The server recognizes specific gestures based on data analysis. For example, drawing a circle with the thumb and index finger is recognized as the "play" gesture.

[0942] Step 10:

[0943] The server interprets and generates a specific command (e.g., play music) based on the recognized gesture.

[0944] Step 11:

[0945] The server sends a feedback message (e.g., "Starting music playback") to the terminal along with the generated command.

[0946] Step 12:

[0947] The device displays a feedback message to the user. For example, it might display "Starting music playback" on the screen.

[0948] Step 13:

[0949] The terminal executes a command. Specifically, it launches the music player application and starts playing music.

[0950] (Example 1)

[0951] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0952] Currently, many users utilize smart devices to access a variety of applications and services. However, these operations are typically performed using touchscreens or physical buttons, often resulting in cumbersome user experiences. Furthermore, traditional methods make intuitive operation and the use of hand gestures difficult, making them particularly inconvenient when hands are dirty or when performing other tasks. A system is needed to solve these problems and provide a more intuitive and convenient way to operate the device.

[0953] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0954] In this invention, the server includes: a shooting device means for acquiring the user's finger movements in real time; an image conversion means for preprocessing image data acquired from the shooting device means; a means for extracting finger feature points from the image data processed by the image conversion means; an analysis device means for analyzing the feature point data; a means for recognizing gestures based on the feature point data analyzed by the analysis device means; a means for interpreting and generating specific commands based on the gesture recognition means; and a terminal device means for executing the commands and feeding the results back to the user. This allows the user to operate intuitively using finger movements, enabling touchless and convenient operation even while working.

[0955] "Photography device means" refers to a device for acquiring the movements of the user's fingers in real time, and includes cameras and the like.

[0956] The "image conversion means" is a means for preprocessing image data acquired from the imaging device means, and performs processing such as grayscale conversion, noise reduction, and contrast adjustment.

[0957] "Methods for extracting feature points of fingers" refer to methods for extracting feature points such as the joints and fingertips of fingers from preprocessed image data, and primarily utilize algorithms and deep learning models.

[0958] "Analysis device means" refers to a means of receiving feature point data transmitted from a terminal and performing analysis using a machine learning model.

[0959] "Means for recognizing gestures" refers to means for classifying and recognizing the type of gesture based on feature point data analyzed by an analysis device.

[0960] "Means for interpreting and generating commands" refers to a method for generating specific commands based on recognized gestures and for interpreting those commands.

[0961] A "terminal device means" is a means of receiving commands sent from a server, actually executing those commands, and providing feedback of the results to the user.

[0962] This invention is a system for recognizing the movements of a user's fingers in real time and performing operations based on those movements. This system consists of the following components: a shooting device, an image conversion device, a device for extracting characteristic points of the fingers, an analysis device, a gesture recognition device, a command interpretation and generation device, and a terminal device.

[0963] System Configuration

[0964] Photography device means

[0965] The camera built into the device captures the user's hand movements in real time. The captured image data is processed in real time.

[0966] Image conversion means

[0967] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. For example, the OpenCV library is used for this preprocessing. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[0968] A method for extracting characteristic points of the fingers.

[0969] Feature points, such as the joints and fingertips of the fingers, are extracted from the preprocessed image data. For this feature point extraction, a model like the MediaPipe Hand is used. The feature points are obtained as numerical data representing the shape and position of the fingers.

[0970] analysis equipment means

[0971] After receiving feature point data sent from the terminal, the server performs analysis using a machine learning model. This machine learning model is, for example, a gesture recognition model trained with TensorFlow. The server uses this model to recognize the type of gesture based on the newly received data.

[0972] Means of recognizing gestures

[0973] The server classifies the gesture type based on the analysis results. Based on this classification, it interprets and generates the relevant commands.

[0974] Means for interpreting and generating commands

[0975] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[0976] Terminal device means

[0977] The device receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, the device launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback."

[0978] Specific example

[0979] When playing music using gestures

[0980] The user makes a "play" gesture (for example, drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera and captures the user's hand movements in real time. The device preprocesses the acquired image by performing grayscale conversion, noise reduction, contrast adjustment, etc. Feature points of the hands are extracted from the preprocessed image and this data is sent to the server. The server analyzes the feature point data using a machine learning model and recognizes that it is a "play" gesture. Based on this result, the server generates a music playback command and sends it to the device. The device receives this command, launches the music player application, starts music playback, and provides feedback to the user saying, "Music playback will begin."

[0981] This system allows users to intuitively perform various operations using only hand gestures. An example of a prompt to be input to the AI ​​model is as follows:

[0982] Example of a prompt

[0983] Prompt: Create a system that recognizes a user's finger gesture for "play" and plays music accordingly. The steps are as follows:

[0984] 1. Use the device's camera to record the movement of your fingers.

[0985] 2. Preprocess the acquired image data (grayscale conversion, noise reduction, contrast adjustment).

[0986] 3. Extract feature points from the preprocessed image.

[0987] 4. Send the extracted feature point data to the server.

[0988] 5. The server analyzes the feature point data and recognizes the gesture.

[0989] 6. Generate a command based on the recognition result and send it to the terminal.

[0990] 7. The terminal receives the command and plays the music.

[0991] This allows users to operate the system with high operability and an intuitive interface.

[0992] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0993] Step 1:

[0994] Camera activation and image capture

[0995] The device activates its built-in camera and captures the user's finger movements in real time. Suppose the user performs a "play" gesture in front of the camera (for example, drawing a circle with their thumb and index finger). The input is the user's finger movements, and the output is real-time image data. This image data is retained for subsequent processing.

[0996] Specific actions:

[0997] The user moves their fingers in front of the device, and the device activates its camera to take a picture.

[0998] Step 2:

[0999] Image preprocessing

[1000] The device performs preprocessing on image data acquired in real time, including grayscale conversion, noise reduction, and contrast adjustment. This improves the quality of the image data and facilitates subsequent processing. The input is the image data from step 1, and the output is the preprocessed image data.

[1001] Specific actions:

[1002] The device converts the image data to grayscale, applies a noise reduction filter, and adjusts the image contrast.

[1003] Step 3:

[1004] Feature point extraction

[1005] The device extracts feature points of the fingers from pre-processed image data. For feature point extraction, for example, the MediaPipe Hand model is used. This model converts the positions of finger joints and fingertips into numerical data. Pre-processed image data is taken as input, and feature point data of the fingers is generated as output.

[1006] Specific actions:

[1007] The device uses the MediaPipe Hand model to calculate the positions of the joints and fingertips of the fingers and obtains numerical data indicating these positions.

[1008] Step 4:

[1009] Sending feature point data to the server

[1010] The device sends feature point data of the fingers to the server. This data is sent to the server via the network as a POST request. The input is feature point data, and the output is a request sent to the server.

[1011] Specific actions:

[1012] The terminal converts the feature point data into JSON format and sends a POST request to the server's API endpoint.

[1013] Step 5:

[1014] Gesture recognition and classification

[1015] The server uses the received feature point data to analyze it with a machine learning model and recognize the type of gesture. The model used is, for example, a gesture recognition model trained with TensorFlow. The input is feature point data, and the output is the gesture recognition result.

[1016] Specific actions:

[1017] The server inputs the received feature point data into a TensorFlow model and performs the analysis. The model recognizes the "play" gesture and generates the recognition result.

[1018] Step 6:

[1019] Command generation and sending

[1020] The server generates a command corresponding to the recognized gesture and sends that command to the terminal. This command could include, for example, an instruction to play music. The input is the gesture recognition result, and the output is the generated command.

[1021] Specific actions:

[1022] The server analyzes the gesture recognition results and generates a corresponding music playback command. The generated command is then sent to the terminal in JSON format.

[1023] Step 7:

[1024] Command execution and feedback

[1025] The terminal receives and executes a command sent from the server. It launches a music player application, starts playing music, and provides feedback to the user saying, "Starting music playback." The input is a command, and the output is music playback and a feedback message.

[1026] Specific actions:

[1027] The terminal analyzes the music playback command, launches the music player, and starts playback. At the same time, a message saying "Music playback will begin" is displayed on the screen to the user.

[1028] (Application Example 1)

[1029] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1030] In modern factories, improving work efficiency and ensuring worker safety are critical challenges. However, conventional robot operating systems often lack intuitive and rapid response capabilities, potentially leading to operational errors and increased worker burden. In particular, the presence of numerous buttons and switches increases the complexity of manual operation and makes operational errors more likely. To address these challenges, a system is needed that more intuitively recognizes worker movements and provides accurate operational feedback.

[1031] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1032] In this invention, the server includes a camera means for acquiring the user's finger movements in real time, an image processing means for preprocessing image data acquired from the camera means, a means for extracting finger feature points from the image data processed by the image processing means, a server means for analyzing the feature point data, a means for recognizing gestures based on the feature point data analyzed by the server means, a means for interpreting and generating specific commands based on the gesture recognition means, a terminal means for executing the commands and feeding the results back to the user, a means for operating the robot in a specific work environment, and a means for starting, stopping, or adjusting work based on gestures using the means. This enables the operator to intuitively and quickly operate the robot using only finger movements, thereby improving work efficiency and ensuring safety.

[1033] A "camera device" is a device used to capture the movements of the user's fingers in real time, and its role is to acquire image data.

[1034] "Image processing means" refers to techniques used to preprocess image data acquired from a camera and improve the quality of the processed image data.

[1035] The "feature point extraction means" is an algorithm for extracting feature points of fingers from pre-processed image data and converting their position and shape into numerical data.

[1036] A "server system" is a computer server that receives feature point data transmitted from a terminal and analyzes the data using a machine learning model.

[1037] "Gesture recognition means" refers to a technology that classifies and recognizes the type of user gesture based on feature point data analyzed by the server means.

[1038] The "command interpretation and generation means" is a means for interpreting and generating a corresponding command in response to a gesture recognized by the gesture recognition means.

[1039] "Terminal means" refers to equipment or devices that execute commands sent from server means and provide feedback of the results to the user.

[1040] "Means for operating robots in a specific work environment" refers to means used to direct specific actions or tasks when operating robots in a factory or work environment.

[1041] "Means for starting, stopping, or adjusting work based on gestures" refers to technology that recognizes the user's hand gestures and starts, stops, or adjusts the robot's movements in response to those gestures.

[1042] This invention is a system for intuitively operating a robot using specific gestures in a factory work environment. Detailed embodiments of this system are described below.

[1043] Key components of the system

[1044] Camera means

[1045] A camera system is used to capture the user's hand movements in real time. The camera, using a webcam or a dedicated high-resolution camera, captures the hand movements and acquires the data.

[1046] Image processing means

[1047] The acquired image data is preprocessed by image processing equipment. Specifically, grayscale conversion, noise reduction, and contrast adjustment are performed using image processing libraries such as OpenCV. This improves the quality of the image data and facilitates subsequent processing by feature point extraction equipment.

[1048] Feature point extraction means

[1049] Feature points, such as the joints and fingertip positions of the fingers, are extracted from the preprocessed image data. Specific algorithms (e.g., HOG or SIFT) are used to convert the shape and position of the fingers into numerical data.

[1050] Server Means

[1051] Feature point data is sent from the terminal to the server. The server analyzes the feature point data using a machine learning model. This machine learning model is built using frameworks such as TensorFlow and trained on a large amount of gesture data.

[1052] Gesture recognition means

[1053] The server classifies the type of gesture based on the analyzed feature point data. For example, if a user makes a "start" gesture, the server will recognize that gesture correctly.

[1054] Command interpretation and generation means

[1055] Based on the gesture recognition results, the server generates corresponding commands. Specifically, commands such as "start," "stop," and "adjust" are generated.

[1056] Terminal means

[1057] The generated command is sent to the terminal, which then executes the command. For example, if the command "start" is received, the robot will begin operation. The terminal also provides feedback to the user, such as "The robot will begin its work."

[1058] System operation example

[1059] For example, recognizing the "start" gesture will cause the conveyor belt in the factory to begin moving. Recognizing the "stop" gesture will cause the conveyor belt to stop. Recognizing the "adjust" gesture will adjust the robot's speed and direction. In this way, users can intuitively and quickly operate robots using only finger movements.

[1060] Example of a prompt

[1061] The following are examples of prompts to input into a generative AI model in a real-world usage scenario:

[1062] When the "start" gesture is recognized, the conveyor belts inside the factory begin to move.

[1063] When the "stop" gesture is recognized, the conveyor belt will stop.

[1064] When it recognizes the "adjust" gesture, it adjusts the robot's movement speed and direction.

[1065] As described above, the present invention improves work efficiency and worker safety within factories.

[1066] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1067] Step 1:

[1068] The terminal activates its camera to capture the user's hand movements in real time. The video data acquired by the camera is used as input and passed to the next process.

[1069] Step 2:

[1070] The terminal preprocesses the video data acquired from the camera using image processing equipment. Specifically, it uses the OpenCV library to convert the video data to grayscale, and then performs noise reduction and contrast adjustment. The processed image data is then output.

[1071] Step 3:

[1072] Based on the pre-processed image data, the terminal uses a feature point extraction mechanism to identify the positions of the joints and fingertips of the fingers. The identified feature point data is output and sent to the server.

[1073] Step 4:

[1074] The server receives feature point data transmitted from the terminal. Using the received feature point data as input, it analyzes the data using a generative AI model. As a result of the analysis, the gesture is identified.

[1075] Step 5:

[1076] The server generates corresponding commands based on the gesture recognition results analyzed by the generation AI model. In this case, commands such as "start," "stop," and "adjust" are generated depending on the recognized gesture. The generated commands are output and sent to the terminal.

[1077] Step 6:

[1078] The terminal receives commands sent from the server. It uses the received commands as input to execute specific robot operations. For example, upon receiving the "start" command, the terminal starts the conveyor belt. The result of this operation is output.

[1079] Step 7:

[1080] The terminal provides feedback to the user regarding the results of the operations performed. Specifically, it displays messages such as "Robot operation will begin" to the user. This allows the user to confirm that the operation was performed successfully.

[1081] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1082] This invention is a system that combines a system for recognizing the user's hand movements in real time and performing operations based on those movements with an emotion engine for recognizing the user's emotions and reflecting them in the feedback. This system consists of a camera, image processing, feature point extraction, server, gesture recognition, command interpretation and generation, terminal, and emotion engine.

[1083] System Configuration

[1084] Camera means

[1085] The camera built into the device captures the user's hand movements and facial expressions in real time. The captured image data is processed in real time.

[1086] Image processing means

[1087] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[1088] Feature point extraction means

[1089] Feature points, such as the joints and fingertips of the fingers, are extracted from the pre-processed image data. This converts the shape and position of the fingers into numerical data.

[1090] Server Means

[1091] After receiving feature point data sent from the terminal, the server performs analysis using a machine learning model. This machine learning model has been pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[1092] Gesture recognition means

[1093] The server classifies the type of gesture based on the feature point data it analyzes. Based on this classification, it interprets and generates the relevant commands.

[1094] Command interpretation and generation means

[1095] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[1096] Emotional Engine

[1097] The server incorporates an emotion engine to analyze the user's facial expressions, voice, or other biometric data. This emotion engine recognizes the user's emotions and reflects them in the interpretation of gesture commands and the feedback provided.

[1098] Terminal means

[1099] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it might respond with something like, "Playing a calming song."

[1100] Specific example

[1101] When playing music using gestures

[1102] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera, captures the user's hand movements in real time, and simultaneously acquires their facial expressions.

[1103] 2. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[1104] 3. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[1105] 4. Based on this result, the server generates a music playback command and sends it to the terminal, along with a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[1106] 5. The device displays this feedback message, launches the music player application, and begins playing music.

[1107] This system allows users to intuitively perform various operations using only hand gestures, providing high operability and a superior user experience. Furthermore, by recognizing user emotions and responding accordingly, it achieves a more human-like interaction.

[1108] The following describes the processing flow.

[1109] Step 1:

[1110] The device activates its built-in camera. The camera begins capturing the user's hand movements and facial expressions in real time.

[1111] Step 2:

[1112] The device acquires captured images frame by frame in real time. The user's hand movements and facial expressions are continuously collected.

[1113] Step 3:

[1114] The device performs preprocessing on the acquired image data. Specifically, it converts the image to grayscale, applies a noise reduction filter, and adjusts the contrast. This improves the accuracy of subsequent processing.

[1115] Step 4:

[1116] The device segments the finger area from the pre-processed image. This process uses color distribution and contour detection algorithms.

[1117] Step 5:

[1118] The device extracts feature points from segmented finger regions. Specifically, it identifies the positions of finger joints and fingertips and records them as numerical data.

[1119] Step 6:

[1120] The device extracts facial feature points to analyze the user's facial expression data. This feature point data is analyzed in real time.

[1121] Step 7:

[1122] The device sends extracted finger and facial feature point data to the server. This transmission is done via the internet or a local network.

[1123] Step 8:

[1124] The server receives feature point data and prepares it for analysis. The received data is then loaded into the system.

[1125] Step 9:

[1126] The server uses machine learning models to analyze the feature point data of the fingers. This involves algorithms such as neural networks and support vector machines.

[1127] Step 10:

[1128] The server then uses an emotion engine to recognize the user's emotional state from facial feature point data. For example, it can determine "joy" from a smile and "anger" from wrinkles between the eyebrows.

[1129] Step 11:

[1130] The server recognizes the type of gesture (e.g., the "play" gesture) based on the analysis results of the hand feature point data.

[1131] Step 12:

[1132] The server interprets and generates specific commands (e.g., play music) based on recognized gestures and the user's emotional state. It also generates appropriate feedback messages tailored to the emotional state.

[1133] Step 13:

[1134] The server sends the generated command and feedback message (e.g., "Starting music playback," "Playing a song to refresh your mood") to the terminal.

[1135] Step 14:

[1136] The device displays a feedback message to the user. For example, it might display "Starting music playback" on the screen.

[1137] Step 15:

[1138] The terminal executes a command. Specifically, it launches a music player application and starts playing music. The type of music played is also adjusted according to the user's emotional state.

[1139] (Example 2)

[1140] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1141] There is a need to develop a system that can accurately recognize a user's hand movements and emotions in real time and perform actions based on that recognition. However, conventional technologies have insufficient accuracy in image processing, feature point extraction, and gesture recognition, and there are no systems that provide feedback that takes the user's emotional state into account. As a result, there are challenges such as a reduced user experience and difficulty in intuitive operation.

[1142] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1143] In this invention, the server includes a shooting device for acquiring the user's finger movements in real time, an image processing device for preprocessing image data acquired from the shooting device, a device for extracting feature points of the fingers from the image data processed by the image processing device, an analysis device for analyzing the feature point data, a device for recognizing gestures based on the feature point data analyzed by the analysis device, a device for interpreting and generating specific instructions based on the gesture recognition device, a terminal device for executing the instructions and providing feedback to the user, an emotion engine for analyzing the user's facial expression data and recognizing their emotional state, and a device for adjusting the feedback content based on the emotional state recognized by the emotion engine. This enables real-time, highly accurate gesture recognition and appropriate feedback that takes the user's emotions into consideration.

[1144] A "user" refers to a person who operates a system.

[1145] "Hand and finger movements" refers to the position and actions of the user's hands and fingers.

[1146] "Real-time" refers to processing or responding instantly without delay or waiting time.

[1147] "Image acquisition device" refers to a device such as a camera used to acquire images.

[1148] "Image data" refers to the digital information of an image acquired by a camera.

[1149] "Preprocessing" refers to the process of modifying image data to make it easier to analyze and extract feature points.

[1150] An "image processing device" refers to a device used for pre-processing image data.

[1151] "Feature points" refer to important points necessary for analysis, such as specific positions or joint points on the fingers.

[1152] "Extraction" refers to the process of extracting feature points from image data.

[1153] An "analysis device" refers to a device that performs analysis based on feature point data.

[1154] A "gesture" refers to a specific action formed by combining hand and finger movements.

[1155] "Recognition" refers to the process of identifying gestures and emotions based on the analysis results.

[1156] "Instructions" refer to actions or commands that correspond to gestures.

[1157] A "terminal device" refers to an electronic device that a user operates or uses to receive feedback.

[1158] "Facial expression data" refers to digital information obtained from the movements and state of a user's face.

[1159] An "emotion engine" refers to a system that analyzes facial expression data, voice, and other information to recognize a user's emotions.

[1160] "Feedback" refers to the responses or information that a system provides to a user.

[1161] This invention relates to a system that recognizes a user's hand movements and emotions in real time and performs operations based on them. The system includes a camera, an image processing device, a feature point extraction device, an analysis device, a gesture recognition device, an instruction generation device, a terminal device, and an emotion engine.

[1162] System Configuration

[1163] Imaging device

[1164] When a user operates the system, a camera on the terminal captures their hand movements and facial expressions in real time. The captured image data is then directly transmitted to an image processing device.

[1165] Image processing device

[1166] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[1167] Feature point extraction device

[1168] Feature points, such as the joints and fingertips of the fingers, are extracted from the pre-processed image data. This converts the shape and position of the fingers into numerical data, which is then transmitted to the analysis device.

[1169] analysis device

[1170] A server receives feature point data sent from a terminal and performs analysis using a machine learning model. This machine learning model is pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[1171] Gesture recognition device

[1172] The analysis device classifies the type of gesture based on the feature point data it analyzes. For example, it might be recognized as a "play" gesture.

[1173] instruction generation device

[1174] The analysis device generates instructions corresponding to the recognized gesture. These instructions can range from launching a specific application, performing a specific action, to displaying specific data.

[1175] Emotional Engine

[1176] The server incorporates an emotion engine to analyze biometric data such as the user's facial expressions and voice. This emotion engine recognizes the user's emotions and reflects them in the interpretation of gesture-based instructions and feedback.

[1177] Terminal device

[1178] The terminal device receives instructions sent from the server and actually executes those instructions. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it will respond with something like, "Playing a calming song."

[1179] Specific example

[1180] When playing music using gestures

[1181] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera.

[1182] 2. The device activates its camera, capturing the user's hand movements in real time, and simultaneously acquiring their facial expressions.

[1183] 3. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[1184] 4. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[1185] 5. Based on this result, the server generates a music playback command and sends it to the terminal, including a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[1186] 6. The device displays this feedback message, launches the music player application, and begins playing music.

[1187] Example of a prompt

[1188] This system recognizes the user's hand movements and emotions in real time and performs actions based solely on that recognition. It captures hand movements and facial expressions with a camera, and then performs image processing and feature point extraction. The obtained data is sent to a server, which analyzes it to recognize gestures and emotions. The server then generates commands based on the recognition results and executes them on the terminal.

[1189] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1190] Step 1:

[1191] The user makes a hand gesture. The user performs a specific gesture (for example, a "play" gesture where the thumb and index finger draw a circle) in front of the device's camera. The camera captures this action.

[1192] Input: Finger movements

[1193] Output: Raw image data

[1194] Specific operation: The camera captures the user's hand movements in real time and generates image data.

[1195] Step 2:

[1196] The device performs image processing. The device performs preprocessing on image data acquired from the camera, such as grayscale conversion, noise reduction, and contrast adjustment. This processing improves the quality of the image data, enabling smoother subsequent feature point extraction.

[1197] Input: Raw image data

[1198] Output: Preprocessed image data

[1199] Specific operation: Converts a color image to grayscale, removes noise, and adjusts contrast.

[1200] Step 3:

[1201] The device extracts feature points. It extracts feature points such as the joints and fingertips of the fingers from the pre-processed image data. This converts the shape and position of the fingers into numerical data.

[1202] Input: Preprocessed image data

[1203] Output: Feature point data (numerical data)

[1204] Specific operation: Calculates the position of the joints and fingertips of the fingers and generates their coordinate data.

[1205] Step 4:

[1206] The terminal sends feature point data to the server. The terminal sends the extracted feature point data to the server.

[1207] Input: Feature point data

[1208] Output: Data transfer to the server

[1209] Specific operation: Send feature point data in JSON format to a specific server address.

[1210] Step 5:

[1211] The server performs gesture and emotion analysis. The server inputs the received feature point data into a machine learning model to recognize the type of gesture. Simultaneously, the emotion engine analyzes the user's facial expression data to recognize the user's emotional state.

[1212] Input: Feature point data, facial expression data

[1213] Output: Gesture recognition results, emotion recognition results

[1214] Specific operation: Feature point data is input into a machine learning model and recognized as a "play" gesture. Simultaneously, the emotion engine determines "joy" from the facial expression.

[1215] Step 6:

[1216] The server generates commands. The server generates appropriate commands based on the analysis of gestures and emotions. For example, it generates a music playback command for a "play" gesture.

[1217] Input: Gesture recognition results, emotion recognition results

[1218] Output: Music playback command

[1219] Specific operation: Combine the gesture recognition results and emotion recognition results to generate a music playback command in JSON format and send it to the terminal.

[1220] Step 7:

[1221] The server generates feedback messages. The server generates feedback messages that are tailored to the user's emotional state. For example, if the user is happy, it might generate a message such as, "We'll play some music to refresh your mood."

[1222] Input: Emotion recognition result

[1223] Output: Feedback message

[1224] Specific operation: Based on the user's emotional state, generate an appropriate feedback message and send it to the device along with instructions.

[1225] Step 8:

[1226] The terminal executes a command and displays feedback. The terminal executes a command sent from the server, for example, launching a music player application and starting music playback. Simultaneously, a feedback message is displayed to the user.

[1227] Input: Music playback command, feedback message

[1228] Output: Music playback, message display

[1229] Specific actions: Launch the music player and start playing music. Simultaneously, display the message "Playing music to refresh your mood" on the screen.

[1230] (Application Example 2)

[1231] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1232] In modern interactive systems, improving user operability and experience is a crucial challenge. However, conventional systems struggle to accurately recognize user hand movements and gestures, resulting in limitations in operation and a lack of features to provide feedback that responds to user emotions. This limits user interaction, and there is a growing demand for more human-like and flexible responses.

[1233] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1234] In this invention, the server includes: video acquisition means for acquiring the movements of the user's fingers in real time; image processing means for preprocessing image data acquired from the video acquisition means; means for extracting feature points of the fingers from the image data processed by the image processing means; analysis means for analyzing the feature point data; means for recognizing gestures based on the feature point data analyzed by the analysis means; means for interpreting and generating a specific command based on the gesture recognition means; execution means for executing the command and providing feedback of the result to the user; and emotion analysis means for analyzing the user's emotions when the command is executed and providing appropriate feedback. This makes it possible to accurately recognize the gestures of the user's fingers and provide appropriate feedback according to the user's emotions.

[1235] "Video acquisition means" refers to a device for acquiring the user's hand movements and facial expressions in real time.

[1236] "Image processing means" refers to means for pre-processing image data obtained from video acquisition means to improve its quality.

[1237] A "feature point extraction method" is a means of extracting feature points, such as the joints and fingertips of the fingers, from pre-processed image data.

[1238] "Analysis means" refers to a computer-based system that analyzes feature point data and recognizes gestures and emotions based on that analysis.

[1239] A "gesture recognition means" is a means of classifying a user's gestures based on feature point data recognized by an analysis means.

[1240] The "command interpretation and generation means" is a means for interpreting and generating commands corresponding to gestures recognized by the gesture recognition means.

[1241] "Execution means" refers to a means of executing commands generated by command interpretation and generation means and providing feedback of the execution results to the user.

[1242] "Emotion analysis methods" are means of recognizing a user's emotions by analyzing their facial expressions and voice data.

[1243] This invention is a system for recognizing the user's hand movements in real time and executing operations based on those movements, and further incorporates an emotion engine that recognizes the user's emotions and reflects them in the feedback. This system consists of video acquisition means, image processing means, feature point extraction means, analysis means, gesture recognition means, command interpretation and generation means, execution means, and emotion analysis means.

[1244] System Configuration

[1245] Video acquisition method

[1246] It includes a camera to capture the user's hand movements and facial expressions in real time. The camera is attached to a device (e.g., a smartphone or PC). This allows the user's hand movements and facial expressions to be collected.

[1247] Image processing means

[1248] It has a function to preprocess image data acquired from video acquisition devices. This preprocessing includes grayscale conversion, noise reduction, and contrast adjustment. The software used is an open-source image processing library (e.g., OpenCV). This preprocessing improves the quality of the image data and makes subsequent feature point extraction easier.

[1249] Feature point extraction means

[1250] Feature points, such as the joints and fingertips of the fingers, are extracted from pre-processed image data. This process converts the shape and position of the fingers into numerical data. The machine learning model used has been pre-trained on a large amount of gesture data.

[1251] Analysis means

[1252] The server receives feature point data sent from the terminal and analyzes this data using a machine learning model. This machine learning model uses a framework such as TensorFlow. The server analyzes the feature point data to recognize the type of gesture and the user's emotional state.

[1253] Gesture recognition means

[1254] The server identifies the type of gesture based on the analyzed feature point data. Based on this classification, it interprets and generates the relevant commands.

[1255] Command interpretation and generation means

[1256] The server generates commands corresponding to the recognized gestures. These commands are wide-ranging and include launching specific applications, performing specific actions, and displaying specific data.

[1257] Execution method

[1258] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it will respond with something like, "Playing a calming song."

[1259] Emotion analysis means

[1260] The device analyzes facial and voice data it collects to recognize the user's emotions. This emotion analysis utilizes libraries such as the EmotionRecognition library. The emotion analysis provides appropriate feedback for gesture interpretation and generated commands.

[1261] Specific example

[1262] When playing music using gestures

[1263] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera, captures the user's hand movements in real time, and simultaneously acquires their facial expressions.

[1264] 2. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[1265] 3. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[1266] 4. Based on this result, the server generates a music playback command and sends it to the terminal, along with a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[1267] 5. The device displays this feedback message, launches the music player application, and begins playing music.

[1268] Example of a prompt

[1269] "Implement an interactive system that recognizes the user's hand movements in real time and provides appropriate feedback based on the user's emotions. This system recognizes hand gestures using image processing and machine learning, and evaluates the user's emotions using an emotion engine."

[1270] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1271] Step 1:

[1272] The user performs a "play" gesture.

[1273] Input: User's hand and finger movements and facial expressions.

[1274] Action: The user makes a circular gesture with their thumb and index finger in front of the device's camera.

[1275] Step 2:

[1276] The device activates its camera and captures the user's hand movements and facial expressions in real time.

[1277] Input: Video data from the camera.

[1278] Output: Video data of hand and finger movements and facial expressions.

[1279] Operation: The camera records the user's gestures and facial expressions in real time.

[1280] Step 3:

[1281] The device performs pre-processing on the captured video data, including grayscale conversion, noise reduction, and contrast adjustment.

[1282] Input: Video data of hand and finger movements and facial expressions.

[1283] Output: Preprocessed image data.

[1284] Operation: Preprocessing is performed on the terminal side using image processing software (e.g., OpenCV) to improve the quality of the video data.

[1285] Step 4:

[1286] Feature points, such as the positions of finger joints and fingertips, are extracted from the pre-processed image data.

[1287] Input: Preprocessed image data.

[1288] Output: Feature point data for fingers.

[1289] Operation: The terminal uses a feature point extraction algorithm to convert the positions of finger joints and fingertips into numerical data.

[1290] Step 5:

[1291] The device sends characteristic point data of the fingers to the server.

[1292] Input: Feature point data for fingers.

[1293] Output: Sending data to the server.

[1294] Operation: The terminal sends the extracted feature point data to the server via the network.

[1295] Step 6:

[1296] The server analyzes the received feature point data using a machine learning model to recognize the type of gesture.

[1297] Input: Feature point data for fingers.

[1298] Output: Gesture recognition result.

[1299] Operation: Feature point data is analyzed using a machine learning framework (e.g., TensorFlow) running on the server side.

[1300] Step 7:

[1301] The server simultaneously analyzes facial expression data using an emotion engine to recognize the user's emotions.

[1302] Input: Facial expression data.

[1303] Output: Emotion recognition result.

[1304] Operation: The system uses an emotion analysis library (e.g., EmotionRecognition) running on the server side to analyze emotions from facial expression data.

[1305] Step 8:

[1306] The server generates a music playback command based on the gesture recognition results and emotion recognition results.

[1307] Input: Gesture recognition results, emotion recognition results.

[1308] Output: Music playback command.

[1309] Operation: The server interprets the user's intent based on the analysis results and generates an appropriate music playback command.

[1310] Step 9:

[1311] The terminal executes the command received from the server and starts playing music.

[1312] Input: Music playback command.

[1313] Output: Music playback begins.

[1314] Operation: The device launches the music player application and begins playing music.

[1315] Step 10:

[1316] The device displays feedback messages tailored to the user's emotional state.

[1317] Input: Emotion recognition result.

[1318] Output: Feedback message.

[1319] Operation: Based on the emotion recognition results, the device displays a feedback message on the screen that is appropriate for the user (e.g., "Playing a song to refresh your mood").

[1320] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1321] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1322] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1323] [Fourth Embodiment]

[1324] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1325] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1326] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1327] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1328] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1329] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1330] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1331] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1332] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1333] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1334] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1335] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1336] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1337] This invention is a system for recognizing the movements of a user's fingers in real time and performing operations based on those movements. This system consists of the following components: a camera, an image processing unit, a feature point extraction unit, a server, a gesture recognition unit, a command interpretation and generation unit, and a terminal unit.

[1338] System Configuration

[1339] Camera means

[1340] The camera built into the device captures the user's hand movements in real time. The captured image data is processed in real time.

[1341] Image processing means

[1342] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[1343] Feature point extraction means

[1344] Feature points, such as the joints and fingertips of the fingers, are extracted from the preprocessed image data. This feature point extraction uses a specific algorithm to convert the shape and position of the fingers into numerical data.

[1345] Server Means

[1346] After receiving feature point data transmitted from the terminal, the server performs analysis using a machine learning model. This machine learning model has been pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[1347] Gesture recognition means

[1348] The server classifies the type of gesture based on the feature point data it analyzes. Based on this classification, it interprets and generates the relevant commands.

[1349] Command interpretation and generation means

[1350] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[1351] Terminal means

[1352] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback."

[1353] Specific example

[1354] When playing music using gestures

[1355] The user makes a "play" gesture (for example, drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera and captures the user's hand movements in real time. The acquired image is pre-processed by the device, including grayscale conversion, noise reduction, and contrast adjustment. Feature points of the hands are extracted from the pre-processed image, and this data is sent to the server.

[1356] The server uses a machine learning model to analyze feature point data and recognizes it as a "play" gesture. Based on this, the server generates a music playback command and sends it to the terminal. The terminal receives this command, launches the music player application, starts playing music, and provides feedback to the user saying, "Music playback will begin."

[1357] This system allows users to intuitively perform various operations using only hand gestures, providing high operability and a superior user experience.

[1358] The following describes the processing flow.

[1359] Step 1:

[1360] The device activates its built-in camera. The camera begins capturing the user's hand movements in real time.

[1361] Step 2:

[1362] The device acquires captured images frame by frame in real time. This allows for the continuous collection of the user's hand movements.

[1363] Step 3:

[1364] The device performs preprocessing on the acquired image data. Specifically, it converts the image to grayscale, applies a noise reduction filter, and adjusts the contrast. This improves the accuracy of subsequent processing.

[1365] Step 4:

[1366] The device segments the finger area from the pre-processed image. This is done using color distribution and contour detection algorithms.

[1367] Step 5:

[1368] The device extracts feature points from segmented finger regions. Specifically, it identifies the positions of finger joints and fingertips and records them as numerical data.

[1369] Step 6:

[1370] The terminal sends the extracted feature point data to the server. This transmission is done via the internet or a local network.

[1371] Step 7:

[1372] The server receives feature point data and prepares it for analysis. The received data is then loaded into the system.

[1373] Step 8:

[1374] The server analyzes the data using machine learning models. This involves algorithms such as neural networks and support vector machines.

[1375] Step 9:

[1376] The server recognizes specific gestures based on data analysis. For example, drawing a circle with the thumb and index finger is recognized as the "play" gesture.

[1377] Step 10:

[1378] The server interprets and generates a specific command (e.g., play music) based on the recognized gesture.

[1379] Step 11:

[1380] The server sends a feedback message (e.g., "Starting music playback") to the terminal along with the generated command.

[1381] Step 12:

[1382] The device displays a feedback message to the user. For example, it might display "Starting music playback" on the screen.

[1383] Step 13:

[1384] The terminal executes a command. Specifically, it launches the music player application and starts playing music.

[1385] (Example 1)

[1386] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1387] Currently, many users utilize smart devices to access a variety of applications and services. However, these operations are typically performed using touchscreens or physical buttons, often resulting in cumbersome user experiences. Furthermore, traditional methods make intuitive operation and the use of hand gestures difficult, making them particularly inconvenient when hands are dirty or when performing other tasks. A system is needed to solve these problems and provide a more intuitive and convenient way to operate the device.

[1388] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1389] In this invention, the server includes: a shooting device means for acquiring the user's finger movements in real time; an image conversion means for preprocessing image data acquired from the shooting device means; a means for extracting finger feature points from the image data processed by the image conversion means; an analysis device means for analyzing the feature point data; a means for recognizing gestures based on the feature point data analyzed by the analysis device means; a means for interpreting and generating specific commands based on the gesture recognition means; and a terminal device means for executing the commands and feeding the results back to the user. This allows the user to operate intuitively using finger movements, enabling touchless and convenient operation even while working.

[1390] "Photography device means" refers to a device for acquiring the movements of the user's fingers in real time, and includes cameras and the like.

[1391] The "image conversion means" is a means for preprocessing image data acquired from the imaging device means, and performs processing such as grayscale conversion, noise reduction, and contrast adjustment.

[1392] "Methods for extracting feature points of fingers" refer to methods for extracting feature points such as the joints and fingertips of fingers from preprocessed image data, and primarily utilize algorithms and deep learning models.

[1393] "Analysis device means" refers to a means of receiving feature point data transmitted from a terminal and performing analysis using a machine learning model.

[1394] "Means for recognizing gestures" refers to means for classifying and recognizing the type of gesture based on feature point data analyzed by an analysis device.

[1395] "Means for interpreting and generating commands" refers to a method for generating specific commands based on recognized gestures and for interpreting those commands.

[1396] A "terminal device means" is a means of receiving commands sent from a server, actually executing those commands, and providing feedback of the results to the user.

[1397] This invention is a system for recognizing the movements of a user's fingers in real time and performing operations based on those movements. This system consists of the following components: a shooting device, an image conversion device, a device for extracting characteristic points of the fingers, an analysis device, a gesture recognition device, a command interpretation and generation device, and a terminal device.

[1398] System Configuration

[1399] Photography device means

[1400] The camera built into the device captures the user's hand movements in real time. The captured image data is processed in real time.

[1401] Image conversion means

[1402] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. For example, the OpenCV library is used for this preprocessing. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[1403] A method for extracting characteristic points of the fingers.

[1404] Feature points, such as the joints and fingertips of the fingers, are extracted from the preprocessed image data. For this feature point extraction, a model like the MediaPipe Hand is used. The feature points are obtained as numerical data representing the shape and position of the fingers.

[1405] analysis equipment means

[1406] After receiving feature point data sent from the terminal, the server performs analysis using a machine learning model. This machine learning model is, for example, a gesture recognition model trained with TensorFlow. The server uses this model to recognize the type of gesture based on the newly received data.

[1407] Means of recognizing gestures

[1408] The server classifies the gesture type based on the analysis results. Based on this classification, it interprets and generates the relevant commands.

[1409] Means for interpreting and generating commands

[1410] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[1411] Terminal device means

[1412] The device receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, the device launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback."

[1413] Specific example

[1414] When playing music using gestures

[1415] The user makes a "play" gesture (for example, drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera and captures the user's hand movements in real time. The device preprocesses the acquired image by performing grayscale conversion, noise reduction, contrast adjustment, etc. Feature points of the hands are extracted from the preprocessed image and this data is sent to the server. The server analyzes the feature point data using a machine learning model and recognizes that it is a "play" gesture. Based on this result, the server generates a music playback command and sends it to the device. The device receives this command, launches the music player application, starts music playback, and provides feedback to the user saying, "Music playback will begin."

[1416] This system allows users to intuitively perform various operations using only hand gestures. An example of a prompt to be input to the AI ​​model is as follows:

[1417] Example of a prompt

[1418] Prompt: Create a system that recognizes a user's finger gesture for "play" and plays music accordingly. The steps are as follows:

[1419] 1. Use the device's camera to record the movement of your fingers.

[1420] 2. Preprocess the acquired image data (grayscale conversion, noise reduction, contrast adjustment).

[1421] 3. Extract feature points from the preprocessed image.

[1422] 4. Send the extracted feature point data to the server.

[1423] 5. The server analyzes the feature point data and recognizes the gesture.

[1424] 6. Generate a command based on the recognition result and send it to the terminal.

[1425] 7. The terminal receives the command and plays the music.

[1426] This allows users to operate the system with high operability and an intuitive interface.

[1427] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1428] Step 1:

[1429] Camera activation and image capture

[1430] The device activates its built-in camera and captures the user's finger movements in real time. Suppose the user performs a "play" gesture in front of the camera (for example, drawing a circle with their thumb and index finger). The input is the user's finger movements, and the output is real-time image data. This image data is retained for subsequent processing.

[1431] Specific actions:

[1432] The user moves their fingers in front of the device, and the device activates its camera to take a picture.

[1433] Step 2:

[1434] Image preprocessing

[1435] The device performs preprocessing on image data acquired in real time, including grayscale conversion, noise reduction, and contrast adjustment. This improves the quality of the image data and facilitates subsequent processing. The input is the image data from step 1, and the output is the preprocessed image data.

[1436] Specific actions:

[1437] The device converts the image data to grayscale, applies a noise reduction filter, and adjusts the image contrast.

[1438] Step 3:

[1439] Feature point extraction

[1440] The device extracts feature points of the fingers from pre-processed image data. For feature point extraction, for example, the MediaPipe Hand model is used. This model converts the positions of finger joints and fingertips into numerical data. Pre-processed image data is taken as input, and feature point data of the fingers is generated as output.

[1441] Specific actions:

[1442] The device uses the MediaPipe Hand model to calculate the positions of the joints and fingertips of the fingers and obtains numerical data indicating these positions.

[1443] Step 4:

[1444] Sending feature point data to the server

[1445] The device sends feature point data of the fingers to the server. This data is sent to the server via the network as a POST request. The input is feature point data, and the output is a request sent to the server.

[1446] Specific actions:

[1447] The terminal converts the feature point data into JSON format and sends a POST request to the server's API endpoint.

[1448] Step 5:

[1449] Gesture recognition and classification

[1450] The server uses the received feature point data to analyze it with a machine learning model and recognize the type of gesture. The model used is, for example, a gesture recognition model trained with TensorFlow. The input is feature point data, and the output is the gesture recognition result.

[1451] Specific actions:

[1452] The server inputs the received feature point data into a TensorFlow model and performs the analysis. The model recognizes the "play" gesture and generates the recognition result.

[1453] Step 6:

[1454] Command generation and sending

[1455] The server generates a command corresponding to the recognized gesture and sends that command to the terminal. This command could include, for example, an instruction to play music. The input is the gesture recognition result, and the output is the generated command.

[1456] Specific actions:

[1457] The server analyzes the gesture recognition results and generates a corresponding music playback command. The generated command is then sent to the terminal in JSON format.

[1458] Step 7:

[1459] Command execution and feedback

[1460] The terminal receives and executes a command sent from the server. It launches a music player application, starts playing music, and provides feedback to the user saying, "Starting music playback." The input is a command, and the output is music playback and a feedback message.

[1461] Specific actions:

[1462] The terminal analyzes the music playback command, launches the music player, and starts playback. At the same time, a message saying "Music playback will begin" is displayed on the screen to the user.

[1463] (Application Example 1)

[1464] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1465] In modern factories, improving work efficiency and ensuring worker safety are critical challenges. However, conventional robot operating systems often lack intuitive and rapid response capabilities, potentially leading to operational errors and increased worker burden. In particular, the presence of numerous buttons and switches increases the complexity of manual operation and makes operational errors more likely. To address these challenges, a system is needed that more intuitively recognizes worker movements and provides accurate operational feedback.

[1466] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1467] In this invention, the server includes a camera means for acquiring the user's finger movements in real time, an image processing means for preprocessing image data acquired from the camera means, a means for extracting finger feature points from the image data processed by the image processing means, a server means for analyzing the feature point data, a means for recognizing gestures based on the feature point data analyzed by the server means, a means for interpreting and generating specific commands based on the gesture recognition means, a terminal means for executing the commands and feeding the results back to the user, a means for operating the robot in a specific work environment, and a means for starting, stopping, or adjusting work based on gestures using the means. This enables the operator to intuitively and quickly operate the robot using only finger movements, thereby improving work efficiency and ensuring safety.

[1468] A "camera device" is a device used to capture the movements of the user's fingers in real time, and its role is to acquire image data.

[1469] "Image processing means" refers to techniques used to preprocess image data acquired from a camera and improve the quality of the processed image data.

[1470] The "feature point extraction means" is an algorithm for extracting feature points of fingers from pre-processed image data and converting their position and shape into numerical data.

[1471] A "server system" is a computer server that receives feature point data transmitted from a terminal and analyzes the data using a machine learning model.

[1472] "Gesture recognition means" refers to a technology that classifies and recognizes the type of user gesture based on feature point data analyzed by the server means.

[1473] The "command interpretation and generation means" is a means for interpreting and generating a corresponding command in response to a gesture recognized by the gesture recognition means.

[1474] "Terminal means" refers to equipment or devices that execute commands sent from server means and provide feedback of the results to the user.

[1475] "Means for operating robots in a specific work environment" refers to means used to direct specific actions or tasks when operating robots in a factory or work environment.

[1476] "Means for starting, stopping, or adjusting work based on gestures" refers to technology that recognizes the user's hand gestures and starts, stops, or adjusts the robot's movements in response to those gestures.

[1477] This invention is a system for intuitively operating a robot using specific gestures in a factory work environment. Detailed embodiments of this system are described below.

[1478] Key components of the system

[1479] Camera means

[1480] A camera system is used to capture the user's hand movements in real time. The camera, using a webcam or a dedicated high-resolution camera, captures the hand movements and acquires the data.

[1481] Image processing means

[1482] The acquired image data is preprocessed by image processing equipment. Specifically, grayscale conversion, noise reduction, and contrast adjustment are performed using image processing libraries such as OpenCV. This improves the quality of the image data and facilitates subsequent processing by feature point extraction equipment.

[1483] Feature point extraction means

[1484] Feature points, such as the joints and fingertip positions of the fingers, are extracted from the preprocessed image data. Specific algorithms (e.g., HOG or SIFT) are used to convert the shape and position of the fingers into numerical data.

[1485] Server Means

[1486] Feature point data is sent from the terminal to the server. The server analyzes the feature point data using a machine learning model. This machine learning model is built using frameworks such as TensorFlow and trained on a large amount of gesture data.

[1487] Gesture recognition means

[1488] The server classifies the type of gesture based on the analyzed feature point data. For example, if a user makes a "start" gesture, the server will recognize that gesture correctly.

[1489] Command interpretation and generation means

[1490] Based on the gesture recognition results, the server generates corresponding commands. Specifically, commands such as "start," "stop," and "adjust" are generated.

[1491] Terminal means

[1492] The generated command is sent to the terminal, which then executes the command. For example, if the command "start" is received, the robot will begin operation. The terminal also provides feedback to the user, such as "The robot will begin its work."

[1493] System operation example

[1494] For example, recognizing the "start" gesture will cause the conveyor belt in the factory to begin moving. Recognizing the "stop" gesture will cause the conveyor belt to stop. Recognizing the "adjust" gesture will adjust the robot's speed and direction. In this way, users can intuitively and quickly operate robots using only finger movements.

[1495] Example of a prompt

[1496] The following are examples of prompts to input into a generative AI model in a real-world usage scenario:

[1497] When the "start" gesture is recognized, the conveyor belts inside the factory begin to move.

[1498] When the "stop" gesture is recognized, the conveyor belt will stop.

[1499] When it recognizes the "adjust" gesture, it adjusts the robot's movement speed and direction.

[1500] As described above, the present invention improves work efficiency and worker safety within factories.

[1501] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1502] Step 1:

[1503] The terminal activates its camera to capture the user's hand movements in real time. The video data acquired by the camera is used as input and passed to the next process.

[1504] Step 2:

[1505] The terminal preprocesses the video data acquired from the camera using image processing equipment. Specifically, it uses the OpenCV library to convert the video data to grayscale, and then performs noise reduction and contrast adjustment. The processed image data is then output.

[1506] Step 3:

[1507] Based on the pre-processed image data, the terminal uses a feature point extraction mechanism to identify the positions of the joints and fingertips of the fingers. The identified feature point data is output and sent to the server.

[1508] Step 4:

[1509] The server receives feature point data transmitted from the terminal. Using the received feature point data as input, it analyzes the data using a generative AI model. As a result of the analysis, the gesture is identified.

[1510] Step 5:

[1511] The server generates corresponding commands based on the gesture recognition results analyzed by the generation AI model. In this case, commands such as "start," "stop," and "adjust" are generated depending on the recognized gesture. The generated commands are output and sent to the terminal.

[1512] Step 6:

[1513] The terminal receives commands sent from the server. It uses the received commands as input to execute specific robot operations. For example, upon receiving the "start" command, the terminal starts the conveyor belt. The result of this operation is output.

[1514] Step 7:

[1515] The terminal provides feedback to the user regarding the results of the operations performed. Specifically, it displays messages such as "Robot operation will begin" to the user. This allows the user to confirm that the operation was performed successfully.

[1516] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1517] This invention is a system that combines a system for recognizing the user's hand movements in real time and performing operations based on those movements with an emotion engine for recognizing the user's emotions and reflecting them in the feedback. This system consists of a camera, image processing, feature point extraction, server, gesture recognition, command interpretation and generation, terminal, and emotion engine.

[1518] System Configuration

[1519] Camera means

[1520] The camera built into the device captures the user's hand movements and facial expressions in real time. The captured image data is processed in real time.

[1521] Image processing means

[1522] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[1523] Feature point extraction means

[1524] Feature points, such as the joints and fingertips of the fingers, are extracted from the pre-processed image data. This converts the shape and position of the fingers into numerical data.

[1525] Server Means

[1526] After receiving feature point data sent from the terminal, the server performs analysis using a machine learning model. This machine learning model has been pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[1527] Gesture recognition means

[1528] The server classifies the type of gesture based on the feature point data it analyzes. Based on this classification, it interprets and generates the relevant commands.

[1529] Command interpretation and generation means

[1530] The server generates a command corresponding to the recognized gesture. This command can do a wide range of things, such as launching a specific application, performing a specific action, or displaying specific data.

[1531] Emotional Engine

[1532] The server incorporates an emotion engine to analyze the user's facial expressions, voice, or other biometric data. This emotion engine recognizes the user's emotions and reflects them in the interpretation of gesture commands and the feedback provided.

[1533] Terminal means

[1534] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it might respond with something like, "Playing a calming song."

[1535] Specific example

[1536] When playing music using gestures

[1537] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera, captures the user's hand movements in real time, and simultaneously acquires their facial expressions.

[1538] 2. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[1539] 3. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[1540] 4. Based on this result, the server generates a music playback command and sends it to the terminal, along with a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[1541] 5. The device displays this feedback message, launches the music player application, and begins playing music.

[1542] This system allows users to intuitively perform various operations using only hand gestures, providing high operability and a superior user experience. Furthermore, by recognizing user emotions and responding accordingly, it achieves a more human-like interaction.

[1543] The following describes the processing flow.

[1544] Step 1:

[1545] The device activates its built-in camera. The camera begins capturing the user's hand movements and facial expressions in real time.

[1546] Step 2:

[1547] The device acquires captured images frame by frame in real time. The user's hand movements and facial expressions are continuously collected.

[1548] Step 3:

[1549] The device performs preprocessing on the acquired image data. Specifically, it converts the image to grayscale, applies a noise reduction filter, and adjusts the contrast. This improves the accuracy of subsequent processing.

[1550] Step 4:

[1551] The device segments the finger area from the pre-processed image. This process uses color distribution and contour detection algorithms.

[1552] Step 5:

[1553] The device extracts feature points from segmented finger regions. Specifically, it identifies the positions of finger joints and fingertips and records them as numerical data.

[1554] Step 6:

[1555] The device extracts facial feature points to analyze the user's facial expression data. This feature point data is analyzed in real time.

[1556] Step 7:

[1557] The device sends extracted finger and facial feature point data to the server. This transmission is done via the internet or a local network.

[1558] Step 8:

[1559] The server receives feature point data and prepares it for analysis. The received data is then loaded into the system.

[1560] Step 9:

[1561] The server uses machine learning models to analyze the feature point data of the fingers. This involves algorithms such as neural networks and support vector machines.

[1562] Step 10:

[1563] The server then uses an emotion engine to recognize the user's emotional state from facial feature point data. For example, it can determine "joy" from a smile and "anger" from wrinkles between the eyebrows.

[1564] Step 11:

[1565] The server recognizes the type of gesture (e.g., the "play" gesture) based on the analysis results of the hand feature point data.

[1566] Step 12:

[1567] The server interprets and generates specific commands (e.g., play music) based on recognized gestures and the user's emotional state. It also generates appropriate feedback messages tailored to the emotional state.

[1568] Step 13:

[1569] The server sends the generated command and feedback message (e.g., "Starting music playback," "Playing a song to refresh your mood") to the terminal.

[1570] Step 14:

[1571] The device displays a feedback message to the user. For example, it might display "Starting music playback" on the screen.

[1572] Step 15:

[1573] The terminal executes a command. Specifically, it launches a music player application and starts playing music. The type of music played is also adjusted according to the user's emotional state.

[1574] (Example 2)

[1575] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1576] There is a need to develop a system that can accurately recognize a user's hand movements and emotions in real time and perform actions based on that recognition. However, conventional technologies have insufficient accuracy in image processing, feature point extraction, and gesture recognition, and there are no systems that provide feedback that takes the user's emotional state into account. As a result, there are challenges such as a reduced user experience and difficulty in intuitive operation.

[1577] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1578] In this invention, the server includes a shooting device for acquiring the user's finger movements in real time, an image processing device for preprocessing image data acquired from the shooting device, a device for extracting feature points of the fingers from the image data processed by the image processing device, an analysis device for analyzing the feature point data, a device for recognizing gestures based on the feature point data analyzed by the analysis device, a device for interpreting and generating specific instructions based on the gesture recognition device, a terminal device for executing the instructions and providing feedback to the user, an emotion engine for analyzing the user's facial expression data and recognizing their emotional state, and a device for adjusting the feedback content based on the emotional state recognized by the emotion engine. This enables real-time, highly accurate gesture recognition and appropriate feedback that takes the user's emotions into consideration.

[1579] A "user" refers to a person who operates a system.

[1580] "Hand and finger movements" refers to the position and actions of the user's hands and fingers.

[1581] "Real-time" refers to processing or responding instantly without delay or waiting time.

[1582] "Image acquisition device" refers to a device such as a camera used to acquire images.

[1583] "Image data" refers to the digital information of an image acquired by a camera.

[1584] "Preprocessing" refers to the process of modifying image data to make it easier to analyze and extract feature points.

[1585] An "image processing device" refers to a device used for pre-processing image data.

[1586] "Feature points" refer to important points necessary for analysis, such as specific positions or joint points on the fingers.

[1587] "Extraction" refers to the process of extracting feature points from image data.

[1588] An "analysis device" refers to a device that performs analysis based on feature point data.

[1589] A "gesture" refers to a specific action formed by combining hand and finger movements.

[1590] "Recognition" refers to the process of identifying gestures and emotions based on the analysis results.

[1591] "Instructions" refer to actions or commands that correspond to gestures.

[1592] A "terminal device" refers to an electronic device that a user operates or uses to receive feedback.

[1593] "Facial expression data" refers to digital information obtained from the movements and state of a user's face.

[1594] An "emotion engine" refers to a system that analyzes facial expression data, voice, and other information to recognize a user's emotions.

[1595] "Feedback" refers to the responses or information that a system provides to a user.

[1596] This invention relates to a system that recognizes a user's hand movements and emotions in real time and performs operations based on them. The system includes a camera, an image processing device, a feature point extraction device, an analysis device, a gesture recognition device, an instruction generation device, a terminal device, and an emotion engine.

[1597] System Configuration

[1598] Imaging device

[1599] When a user operates the system, a camera on the terminal captures their hand movements and facial expressions in real time. The captured image data is then directly transmitted to an image processing device.

[1600] Image processing device

[1601] The device performs preprocessing on image data acquired from the camera, including grayscale conversion, noise reduction, and contrast adjustment. This preprocessing improves the quality of the image data and facilitates the subsequent feature point extraction process.

[1602] Feature point extraction device

[1603] Feature points, such as the joints and fingertips of the fingers, are extracted from the pre-processed image data. This converts the shape and position of the fingers into numerical data, which is then transmitted to the analysis device.

[1604] analysis device

[1605] A server receives feature point data sent from a terminal and performs analysis using a machine learning model. This machine learning model is pre-trained with a large amount of gesture data and recognizes the type of gesture based on the newly received data.

[1606] Gesture recognition device

[1607] The analysis device classifies the type of gesture based on the feature point data it analyzes. For example, it might be recognized as a "play" gesture.

[1608] instruction generation device

[1609] The analysis device generates instructions corresponding to the recognized gesture. These instructions can range from launching a specific application, performing a specific action, to displaying specific data.

[1610] Emotional Engine

[1611] The server incorporates an emotion engine to analyze biometric data such as the user's facial expressions and voice. This emotion engine recognizes the user's emotions and reflects them in the interpretation of gesture-based instructions and feedback.

[1612] Terminal device

[1613] The terminal device receives instructions sent from the server and actually executes those instructions. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it will respond with something like, "Playing a calming song."

[1614] Specific example

[1615] When playing music using gestures

[1616] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera.

[1617] 2. The device activates its camera, capturing the user's hand movements in real time, and simultaneously acquiring their facial expressions.

[1618] 3. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[1619] 4. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[1620] 5. Based on this result, the server generates a music playback command and sends it to the terminal, including a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[1621] 6. The device displays this feedback message, launches the music player application, and begins playing music.

[1622] Example of a prompt

[1623] This system recognizes the user's hand movements and emotions in real time and performs actions based solely on that recognition. It captures hand movements and facial expressions with a camera, and then performs image processing and feature point extraction. The obtained data is sent to a server, which analyzes it to recognize gestures and emotions. The server then generates commands based on the recognition results and executes them on the terminal.

[1624] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1625] Step 1:

[1626] The user makes a hand gesture. The user performs a specific gesture (for example, a "play" gesture where the thumb and index finger draw a circle) in front of the device's camera. The camera captures this action.

[1627] Input: Finger movements

[1628] Output: Raw image data

[1629] Specific operation: The camera captures the user's hand movements in real time and generates image data.

[1630] Step 2:

[1631] The device performs image processing. The device performs preprocessing on image data acquired from the camera, such as grayscale conversion, noise reduction, and contrast adjustment. This processing improves the quality of the image data, enabling smoother subsequent feature point extraction.

[1632] Input: Raw image data

[1633] Output: Preprocessed image data

[1634] Specific operation: Converts a color image to grayscale, removes noise, and adjusts contrast.

[1635] Step 3:

[1636] The device extracts feature points. It extracts feature points such as the joints and fingertips of the fingers from the pre-processed image data. This converts the shape and position of the fingers into numerical data.

[1637] Input: Preprocessed image data

[1638] Output: Feature point data (numerical data)

[1639] Specific operation: Calculates the position of the joints and fingertips of the fingers and generates their coordinate data.

[1640] Step 4:

[1641] The terminal sends feature point data to the server. The terminal sends the extracted feature point data to the server.

[1642] Input: Feature point data

[1643] Output: Data transfer to the server

[1644] Specific operation: Send feature point data in JSON format to a specific server address.

[1645] Step 5:

[1646] The server performs gesture and emotion analysis. The server inputs the received feature point data into a machine learning model to recognize the type of gesture. Simultaneously, the emotion engine analyzes the user's facial expression data to recognize the user's emotional state.

[1647] Input: Feature point data, facial expression data

[1648] Output: Gesture recognition results, emotion recognition results

[1649] Specific operation: Feature point data is input into a machine learning model and recognized as a "play" gesture. Simultaneously, the emotion engine determines "joy" from the facial expression.

[1650] Step 6:

[1651] The server generates commands. The server generates appropriate commands based on the analysis of gestures and emotions. For example, it generates a music playback command for a "play" gesture.

[1652] Input: Gesture recognition results, emotion recognition results

[1653] Output: Music playback command

[1654] Specific operation: Combine the gesture recognition results and emotion recognition results to generate a music playback command in JSON format and send it to the terminal.

[1655] Step 7:

[1656] The server generates feedback messages. The server generates feedback messages that are tailored to the user's emotional state. For example, if the user is happy, it might generate a message such as, "We'll play some music to refresh your mood."

[1657] Input: Emotion recognition result

[1658] Output: Feedback message

[1659] Specific operation: Based on the user's emotional state, generate an appropriate feedback message and send it to the device along with instructions.

[1660] Step 8:

[1661] The terminal executes a command and displays feedback. The terminal executes a command sent from the server, for example, launching a music player application and starting music playback. Simultaneously, a feedback message is displayed to the user.

[1662] Input: Music playback command, feedback message

[1663] Output: Music playback, message display

[1664] Specific actions: Launch the music player and start playing music. Simultaneously, display the message "Playing music to refresh your mood" on the screen.

[1665] (Application Example 2)

[1666] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1667] In modern interactive systems, improving user operability and experience is a crucial challenge. However, conventional systems struggle to accurately recognize user hand movements and gestures, resulting in limitations in operation and a lack of features to provide feedback that responds to user emotions. This limits user interaction, and there is a growing demand for more human-like and flexible responses.

[1668] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1669] In this invention, the server includes: video acquisition means for acquiring the movements of the user's fingers in real time; image processing means for preprocessing image data acquired from the video acquisition means; means for extracting feature points of the fingers from the image data processed by the image processing means; analysis means for analyzing the feature point data; means for recognizing gestures based on the feature point data analyzed by the analysis means; means for interpreting and generating a specific command based on the gesture recognition means; execution means for executing the command and providing feedback of the result to the user; and emotion analysis means for analyzing the user's emotions when the command is executed and providing appropriate feedback. This makes it possible to accurately recognize the gestures of the user's fingers and provide appropriate feedback according to the user's emotions.

[1670] "Video acquisition means" refers to a device for acquiring the user's hand movements and facial expressions in real time.

[1671] "Image processing means" refers to means for pre-processing image data obtained from video acquisition means to improve its quality.

[1672] A "feature point extraction method" is a means of extracting feature points, such as the joints and fingertips of the fingers, from pre-processed image data.

[1673] "Analysis means" refers to a computer-based system that analyzes feature point data and recognizes gestures and emotions based on that analysis.

[1674] A "gesture recognition means" is a means of classifying a user's gestures based on feature point data recognized by an analysis means.

[1675] The "command interpretation and generation means" is a means for interpreting and generating commands corresponding to gestures recognized by the gesture recognition means.

[1676] "Execution means" refers to a means of executing commands generated by command interpretation and generation means and providing feedback of the execution results to the user.

[1677] "Emotion analysis methods" are means of recognizing a user's emotions by analyzing their facial expressions and voice data.

[1678] This invention is a system for recognizing the user's hand movements in real time and executing operations based on those movements, and further incorporates an emotion engine that recognizes the user's emotions and reflects them in the feedback. This system consists of video acquisition means, image processing means, feature point extraction means, analysis means, gesture recognition means, command interpretation and generation means, execution means, and emotion analysis means.

[1679] System Configuration

[1680] Video acquisition method

[1681] It includes a camera to capture the user's hand movements and facial expressions in real time. The camera is attached to a device (e.g., a smartphone or PC). This allows the user's hand movements and facial expressions to be collected.

[1682] Image processing means

[1683] It has a function to preprocess image data acquired from video acquisition devices. This preprocessing includes grayscale conversion, noise reduction, and contrast adjustment. The software used is an open-source image processing library (e.g., OpenCV). This preprocessing improves the quality of the image data and makes subsequent feature point extraction easier.

[1684] Feature point extraction means

[1685] Feature points, such as the joints and fingertips of the fingers, are extracted from pre-processed image data. This process converts the shape and position of the fingers into numerical data. The machine learning model used has been pre-trained on a large amount of gesture data.

[1686] Analysis means

[1687] The server receives feature point data sent from the terminal and analyzes this data using a machine learning model. This machine learning model uses a framework such as TensorFlow. The server analyzes the feature point data to recognize the type of gesture and the user's emotional state.

[1688] Gesture recognition means

[1689] The server identifies the type of gesture based on the analyzed feature point data. Based on this classification, it interprets and generates the relevant commands.

[1690] Command interpretation and generation means

[1691] The server generates commands corresponding to the recognized gestures. These commands are wide-ranging and include launching specific applications, performing specific actions, and displaying specific data.

[1692] Execution method

[1693] The terminal receives commands sent from the server and actually executes them. For example, if a gesture for playing music is recognized, it launches the music player application and starts playback. It also displays a message to the user as feedback, such as "Starting music playback." Furthermore, the feedback is adjusted based on the user's emotional state. For example, if the user is angry, it will respond with something like, "Playing a calming song."

[1694] Emotion analysis means

[1695] The device analyzes facial and voice data it collects to recognize the user's emotions. This emotion analysis utilizes libraries such as the EmotionRecognition library. The emotion analysis provides appropriate feedback for gesture interpretation and generated commands.

[1696] Specific example

[1697] When playing music using gestures

[1698] 1. The user makes a "play" gesture (drawing a circle with their thumb and index finger) in front of the device's camera. The device activates the camera, captures the user's hand movements in real time, and simultaneously acquires their facial expressions.

[1699] 2. The acquired images undergo preprocessing by the terminal, such as grayscale conversion, noise reduction, and contrast adjustment. Feature points of the fingers are extracted from the preprocessed images, and this data is sent to the server.

[1700] 3. The server uses a machine learning model to analyze feature point data and recognizes that it is a "play" gesture. Simultaneously, the emotion engine analyzes the user's facial expression data and recognizes the user's emotional state (e.g., joy, anger, sadness, etc.).

[1701] 4. Based on this result, the server generates a music playback command and sends it to the terminal, along with a feedback message that corresponds to the user's emotional state (e.g., "Starting music playback," "Playing a song to refresh your mood").

[1702] 5. The device displays this feedback message, launches the music player application, and begins playing music.

[1703] Example of a prompt

[1704] "Implement an interactive system that recognizes the user's hand movements in real time and provides appropriate feedback based on the user's emotions. This system recognizes hand gestures using image processing and machine learning, and evaluates the user's emotions using an emotion engine."

[1705] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1706] Step 1:

[1707] The user performs a "play" gesture.

[1708] Input: User's hand and finger movements and facial expressions.

[1709] Action: The user makes a circular gesture with their thumb and index finger in front of the device's camera.

[1710] Step 2:

[1711] The device activates its camera and captures the user's hand movements and facial expressions in real time.

[1712] Input: Video data from the camera.

[1713] Output: Video data of hand and finger movements and facial expressions.

[1714] Operation: The camera records the user's gestures and facial expressions in real time.

[1715] Step 3:

[1716] The device performs pre-processing on the captured video data, including grayscale conversion, noise reduction, and contrast adjustment.

[1717] Input: Video data of hand and finger movements and facial expressions.

[1718] Output: Preprocessed image data.

[1719] Operation: Preprocessing is performed on the terminal side using image processing software (e.g., OpenCV) to improve the quality of the video data.

[1720] Step 4:

[1721] Feature points, such as the positions of finger joints and fingertips, are extracted from the pre-processed image data.

[1722] Input: Preprocessed image data.

[1723] Output: Feature point data for fingers.

[1724] Operation: The terminal uses a feature point extraction algorithm to convert the positions of finger joints and fingertips into numerical data.

[1725] Step 5:

[1726] The device sends characteristic point data of the fingers to the server.

[1727] Input: Feature point data for fingers.

[1728] Output: Sending data to the server.

[1729] Operation: The terminal sends the extracted feature point data to the server via the network.

[1730] Step 6:

[1731] The server analyzes the received feature point data using a machine learning model to recognize the type of gesture.

[1732] Input: Feature point data for fingers.

[1733] Output: Gesture recognition result.

[1734] Operation: Feature point data is analyzed using a machine learning framework (e.g., TensorFlow) running on the server side.

[1735] Step 7:

[1736] The server simultaneously analyzes facial expression data using an emotion engine to recognize the user's emotions.

[1737] Input: Facial expression data.

[1738] Output: Emotion recognition result.

[1739] Operation: The system uses an emotion analysis library (e.g., EmotionRecognition) running on the server side to analyze emotions from facial expression data.

[1740] Step 8:

[1741] The server generates a music playback command based on the gesture recognition results and emotion recognition results.

[1742] Input: Gesture recognition results, emotion recognition results.

[1743] Output: Music playback command.

[1744] Operation: The server interprets the user's intent based on the analysis results and generates an appropriate music playback command.

[1745] Step 9:

[1746] The terminal executes the command received from the server and starts playing music.

[1747] Input: Music playback command.

[1748] Output: Music playback begins.

[1749] Operation: The device launches the music player application and begins playing music.

[1750] Step 10:

[1751] The device displays feedback messages tailored to the user's emotional state.

[1752] Input: Emotion recognition result.

[1753] Output: Feedback message.

[1754] Operation: Based on the emotion recognition results, the device displays a feedback message on the screen that is appropriate for the user (e.g., "Playing a song to refresh your mood").

[1755] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1756] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1757] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1758] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1759] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1760] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1761] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1762] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1763] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1764] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1765] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1766] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1767] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1768] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1769] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1770] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1771] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1772] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1773] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1774] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1775] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1776] The following is further disclosed regarding the embodiments described above.

[1777] (Claim 1)

[1778] A camera system for acquiring the user's hand movements in real time,

[1779] Image processing means for preprocessing image data acquired from the camera means,

[1780] A means for extracting characteristic points of fingers from image data processed by the aforementioned image processing means,

[1781] A server means for analyzing the aforementioned feature point data,

[1782] A means for recognizing gestures based on feature point data analyzed by the server means,

[1783] Means for interpreting and generating a specific command based on the gesture recognition means,

[1784] A terminal means that executes the aforementioned command and provides feedback of the result to the user,

[1785] A system that includes this.

[1786] (Claim 2)

[1787] The system according to claim 1, further comprising means for pre-processing image data collected by the camera means, such as grayscale conversion, noise reduction, and contrast adjustment.

[1788] (Claim 3)

[1789] The system according to claim 1, wherein the server means includes means for analyzing feature point data using a machine learning model and recognizing the type of gesture.

[1790] "Example 1"

[1791] (Claim 1)

[1792] A camera device for acquiring the user's finger movements in real time,

[1793] Image conversion means for preprocessing image data acquired from the aforementioned imaging device means,

[1794] A means for extracting feature points of fingers from image data processed by the aforementioned image conversion means,

[1795] An analysis apparatus means for analyzing the aforementioned feature point data,

[1796] A means for recognizing gestures based on feature point data analyzed by the aforementioned analysis device means,

[1797] Means for interpreting and generating a specific command based on the gesture recognition means,

[1798] A terminal device means that executes the aforementioned command and provides feedback of the result to the user,

[1799] A system that includes this.

[1800] (Claim 2)

[1801] The system according to claim 1, further comprising means for pre-processing the image data collected by the aforementioned imaging device means, such as grayscale conversion, noise reduction, and contrast adjustment.

[1802] (Claim 3)

[1803] The system according to claim 1, wherein the analysis device means includes means for analyzing feature point data using a machine learning model and recognizing the type of gesture.

[1804] "Application Example 1"

[1805] (Claim 1)

[1806] A camera system for acquiring the user's hand movements in real time,

[1807] Image processing means for preprocessing image data acquired from the camera means,

[1808] A means for extracting characteristic points of fingers from image data processed by the aforementioned image processing means,

[1809] A server means for analyzing the aforementioned feature point data,

[1810] A means for recognizing gestures based on feature point data analyzed by the server means,

[1811] Means for interpreting and generating a specific command based on the gesture recognition means,

[1812] A terminal means that executes the aforementioned command and provides feedback of the result to the user,

[1813] Means for operating robots in a specific work environment,

[1814] Means for starting, stopping, or adjusting work based on gestures using the aforementioned means,

[1815] A system that includes this.

[1816] (Claim 2)

[1817] The system according to claim 1, further comprising means for pre-processing image data collected by the camera means, such as grayscale conversion, noise reduction, and contrast adjustment.

[1818] (Claim 3)

[1819] The system according to claim 1, wherein the server means includes means for analyzing feature point data using a machine learning model and recognizing the type of gesture.

[1820] "Example 2 of combining an emotion engine"

[1821] (Claim 1)

[1822] A camera for capturing the user's finger movements in real time,

[1823] An image processing device for preprocessing image data acquired from the aforementioned imaging device,

[1824] A device for extracting characteristic points of fingers from image data processed by the aforementioned image processing device,

[1825] An analysis device for analyzing the aforementioned feature point data,

[1826] A device that recognizes gestures based on feature point data analyzed by the aforementioned analysis device,

[1827] A device that interprets and generates specific instructions based on the gesture recognition device,

[1828] A terminal device that executes the aforementioned instructions and provides feedback of the results to the user,

[1829] An emotion engine that analyzes the user's facial expression data and recognizes their emotional state,

[1830] A device that adjusts the feedback content based on the emotional state recognized by the emotion engine,

[1831] A system that includes this.

[1832] (Claim 2)

[1833] The system according to claim 1, further comprising a device that performs preprocessing on image data collected by the aforementioned imaging device, such as grayscale conversion, noise reduction, and contrast adjustment.

[1834] (Claim 3)

[1835] The system according to claim 1, which includes an analysis device that uses a machine learning model to analyze feature point data and recognize the type of gesture.

[1836] "Application example 2 when combining with an emotional engine"

[1837] (Claim 1)

[1838] A means for acquiring video data to capture the user's finger movements in real time,

[1839] Image processing means for preprocessing image data acquired from the aforementioned video acquisition means,

[1840] A means for extracting characteristic points of fingers from image data processed by the aforementioned image processing means,

[1841] An analysis means for analyzing the aforementioned feature point data,

[1842] A means for recognizing a gesture based on the feature point data analyzed by the aforementioned analysis means,

[1843] Means for interpreting and generating a specific command based on the gesture recognition means,

[1844] An execution means that executes the aforementioned command and provides feedback of the result to the user,

[1845] An emotion analysis means that analyzes the user's emotions when the command is executed and provides appropriate feedback,

[1846] A system that includes this.

[1847] (Claim 2)

[1848] The system according to claim 1, further comprising means for pre-processing the image data collected by the video acquisition means, such as grayscale conversion, noise reduction, and contrast adjustment.

[1849] (Claim 3)

[1850] The system according to claim 1, which includes means for analyzing feature point data using a machine learning model and recognizing the type of gesture. [Explanation of Symbols]

[1851] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A camera system for acquiring the user's hand movements in real time, Image processing means for preprocessing image data acquired from the camera means, A means for extracting characteristic points of fingers from image data processed by the aforementioned image processing means, A server means for analyzing the aforementioned feature point data, A means for recognizing gestures based on feature point data analyzed by the server means, Means for interpreting and generating a specific command based on the gesture recognition means, A terminal means that executes the aforementioned command and provides feedback of the result to the user, A system that includes this.

2. The system according to claim 1, further comprising means for pre-processing the image data collected by the camera means, such as grayscale conversion, noise reduction, and contrast adjustment.

3. The system according to claim 1, wherein the server means includes means for analyzing feature point data using a machine learning model and recognizing the type of gesture.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A