Method and apparatus for human-robot interaction using neural network
The neural network-based human-robot interaction device integrates object detection, face keypoint estimation, and pose detection to address the challenge of real-time mobile operation in diverse environments, enabling accurate and intuitive interaction by interpreting human intentions and emotions.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Patents
- Current Assignee / Owner
- KOREA ELECTRONICS TECH INST
- Filing Date
- 2023-12-26
- Publication Date
- 2026-07-29
AI Technical Summary
Conventional robotic technologies lack optimization for real-time mobile operation in diverse and unpredictable human-robot interaction environments, failing to effectively interpret human intentions and emotions in real-world settings.
A neural network-based human-robot interaction device and method integrating object detection, face keypoint estimation, and pose detection to recognize human behaviors, postures, and emotions, enabling real-time and flexible interaction.
Enables accurate and intuitive human-robot interaction in various environments by comprehensively detecting and interpreting visual cues, allowing robots to respond sensitively to human emotions and intentions, and operate smoothly on mobile devices.
Smart Images

Figure 112023145847884-PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a real-time visual recognition-based human-robot interaction technology for enhancing interaction between a robot and a human. Background Technology
[0002] As robotic technology continues to advance, integrating robots into daily human life is becoming increasingly important. At the heart of this integration lies the challenge of facilitating seamless interaction between humans and robots. Traditionally, interaction between robots and humans has been implemented in limited environments. Most robots are designed to operate within fixed standards and predefined settings, and the utilization of visual recognition modules for human interaction has been restricted.
[0003] While robots were once primarily used in industrial settings, their applications are now expanding into our daily lives. This change was made possible thanks to the powerful performance of deep neural network-based visual recognition modules that surpass existing methods.
[0004] However, while conventional technology addressed the importance of visual perception in robotics, it lacked optimization for real-time mobile operation, resulting in limitations in human-robot interaction in real-world environments. Robots such as delivery and service robots must no longer operate within fixed standards and predefined environments like industrial robots, but must perform tasks flawlessly in diverse and unpredictable settings. In such environments, one of the most significant variables is the human element. Therefore, modern intelligent robots must be developed with human-robot interaction in mind and be capable of discerning human intentions. Prior art literature
[0005] Sheridan, Thomas B. “Human-robot interaction: status and challenges.” Human factors 58.4 (2016): 525-532.Hancock, Peter A., et al. “A meta-analysis of factors affecting trust in human-robot interaction.” Human factors 53.5 (2011): 517-527.Dautenhahn, Kerstin. “Socially intelligent robots: dimensions of human-robot interaction.” Philosophical transactions of the royal society B: Biological sciences 362.1480 (2007): 679-704.Steinfeld, Aaron, et al. “Common metrics for human-robot interaction.” Proceedings of the 1st ACM SIGCHI / SIGART conference on Humanrobot interaction. 2006. De Santis, Agostino, et al. “An atlas of physical human-robot interaction.” Mechanism and Machine Theory 43.3 (2008): 253-270. The problem to be solved
[0006] The present invention aims to provide a human-robot interaction method and device using a neural network, designed to detect and interpret human visual cues in real time to understand and respond to human intentions and emotions.
[0007] Specifically, the present invention aims to provide a human-robot interaction method and device using a neural network that enables more intuitive interaction by integrating object detection, face keypoint estimation, and pose detection functions.
[0008] The objectives of the present invention are not limited to those mentioned above, and other unmentioned objectives will be clearly understood by those skilled in the art from the description below. means of solving the problem
[0009] A human-robot interaction device using a neural network according to one embodiment of the present invention includes: a behavior recognition unit that generates information regarding one or more objects around a subject and information regarding the subject's posture using a neural network-based behavior recognition model based on an input image, and recognizes the subject's behavior based on the information regarding the objects and the information regarding the posture; and a service providing unit that determines a service to be provided to the subject in response to the subject's behavior.
[0010] In one embodiment of the present invention, information regarding the object includes the recognition result of the object and the location of the object. The action recognition unit recognizes the object and determines the location of the object using an RCNN-based object detection model included in the action recognition model based on the input image.
[0011] In one embodiment of the present invention, information regarding the posture includes the posture of the subject and the direction of the subject's movement. The behavior recognition unit recognizes a plurality of key points on the subject's body using a posture estimation model included in the behavior recognition model based on the input image, and determines the subject's posture and the direction of the subject's movement based on the plurality of body key points.
[0012] In one embodiment of the present invention, the behavior recognition unit further generates information regarding the face of the subject using the behavior recognition model based on the input image, and recognizes the behavior of the subject based on information regarding the object, information regarding the posture, and information regarding the face.
[0013] In one embodiment of the present invention, the service providing unit determines whether a service is needed for the subject based on information regarding the posture, and if it is determined that a service is needed for the subject, determines a service to be provided to the subject in response to the subject's actions.
[0014] In one embodiment of the present invention, the information regarding the face includes the face direction, gaze, and emotion of the subject. The behavior recognition unit generates coordinates of the subject's face features using a face feature recognition model included in the behavior recognition model based on the input image, and determines the subject's face direction, gaze, and emotion based on the coordinates of the face features.
[0015] A human-robot interaction method using a neural network according to one embodiment of the present invention comprises: a step in which a human-robot interaction device generates information regarding one or more objects around a subject and information regarding the subject's posture using a neural network-based behavior recognition model based on an input image, and recognizes the subject's behavior based on the information regarding the objects and the information regarding the posture; and a step in which the human-robot interaction device determines a service to be provided to the subject in response to the subject's behavior.
[0016] In one embodiment of the present invention, the human-robot interaction method further includes the step of the human-robot interaction device determining whether a service is required for the subject based on information regarding the posture. The step of determining the service to be provided includes, if the human-robot interaction device determines that a service is required for the subject, determining the service to be provided to the subject in response to the subject's behavior. Effects of the invention
[0017] According to one embodiment of the present invention, the following effects can be obtained.
[0018] (1) Integrated visual recognition: Through the device and method according to the present invention, object detection, face keypoint estimation, and pose detection can be integrated to comprehensively detect and interpret various visual cues of humans. This integrated approach enables the robot to more accurately grasp human intentions and emotions.
[0019] (2) Real-time mobile optimization: The device and method according to the present invention are specifically optimized to operate smoothly on major mobile devices. This allows the robot to perform real-time interaction with humans in a real environment, thereby providing more natural interaction.
[0020] (3) Flexibility in various environments: Through the device and method according to the present invention, interaction with humans can be performed smoothly even in diverse and unpredictable environments. This means that the robot can be effectively utilized in various scenarios in daily life.
[0021] (4) Deep understanding of emotions and intentions: Through the device and method according to the present invention, emotions and intentions can be interpreted quickly and accurately through human facial expressions and gestures. Through this, the robot can respond sensitively to human emotions and take appropriate actions.
[0022] (5) Scalability and Scope of Application: The device and method according to the present invention are designed to be easily applied to various robot platforms and mobile devices. This expands the scope of application of robot technology and enables human-robot interaction in various situations.
[0023] The effects obtainable from the present invention are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art from the description below. Brief explanation of the drawing
[0024] FIG. 1 is a block diagram showing the configuration of a human-robot interaction device according to one embodiment of the present invention. FIG. 2 is a flowchart illustrating a human-robot interaction method according to an embodiment of the present invention. FIG. 3 is a screen showing an experiment on behavior recognition of a human-robot interaction device and method according to an embodiment of the present invention. FIG. 4 is a block diagram showing a computer system for implementing a human-robot interaction method according to an embodiment of the present invention. Specific details for implementing the invention
[0025] The references to the present invention are as [1] below. In this specification, the references or methodologies proposed in the references may be referred to by the numbers assigned to the references below. The full contents of reference [1] are incorporated into this specification.
[0026] [1] Taehyeon Kim, Seho Park, Junho Kwon. “Mobilizing Visual Perception: A Strategy for Enhanced Human-Robot Interaction.” Proceedings of the 2023 International Conference on Communication and Computer Research. 2023.
[0027] This invention is designed to detect and interpret human visual cues in real time, enabling the robot to understand and respond to human intentions and emotions. Furthermore, it is optimized for effective operation on major mobile devices and integrates object detection, facial keypoint estimation, and pose detection capabilities to provide much more intuitive human-robot interaction.
[0028] The present invention is not intended to detect only specific behaviors of a subject viewed from a specific angle of view. The present invention is intended to acquire and analyze information regarding a subject from a general perspective and to provide appropriate services according to the subject's state or behavior.
[0029] The human-robot interaction device according to the present invention can be mounted on a mobile device including a robot. Therefore, the device and method according to the present invention do not require the subject to necessarily enter the field of view, and there is no problem in providing services even if the subject is located at various positions within the field of view. Since the human-robot interaction device according to the present invention includes a posture detection module, an object detection module, or a face keypoint estimation module, it can recognize human behavior more accurately, and accordingly, provide services suitable for the recognized behavior of the subject. For example, even if a specific joint of the subject does not necessarily enter the camera detection area, the device according to the present invention can infer the subject's behavior based on complex information by utilizing the object detection module, thereby enabling accurate behavior recognition.
[0030] Furthermore, the device and method according to the present invention can utilize results perceived through multiple cameras as well as a single movable small robot. That is, various information that allows for the intuitive inference of user behavior can be acquired by utilizing cameras. Therefore, through the device and method according to the present invention, services can be provided to a subject proactively within established norms, rather than simply providing information of interest to the subject and waiting for instructions from the subject.
[0031] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims. Meanwhile, the terms used in this specification are for describing the embodiments and are not intended to limit the present invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. The terms "comprises" and / or "comprising" as used in this specification do not exclude the presence or addition of one or more other components, steps, actions, and / or elements in addition to the mentioned components, steps, actions, and / or elements.
[0032] In describing the present invention, detailed descriptions of related prior art are omitted if it is determined that such descriptions may unnecessarily obscure the essence of the invention.
[0033] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. In order to facilitate an overall understanding in describing the present invention, the same reference numerals will be used for the same means regardless of the drawing number.
[0034] FIG. 1 is a block diagram showing the configuration of a human-robot interaction device according to one embodiment of the present invention.
[0035] A human-robot interaction device (100, hereinafter abbreviated as 'human-robot interaction device') using a neural network according to one embodiment of the present invention includes an action recognition unit (110) and a service provision unit (120), and may further include a speech unit (130), a voice recognition unit (140), and a model training unit (150). The human-robot interaction device (100) illustrated in FIG. 1 is according to one embodiment, and the components of the human-robot interaction device (100) according to the present invention are not limited to the embodiment illustrated in FIG. 1 and may be added, changed, or deleted as necessary.
[0036] The human-robot interaction device (100) is a device that recognizes the actions of a subject (human) and, in response to said actions, provides services for the subject when it is determined that such actions are necessary.
[0037] The behavior recognition unit (110) generates information regarding one or more objects surrounding the subject, information regarding the subject's posture, and information regarding the subject's face, or a combination thereof, using a neural network-based behavior recognition model based on an input image (e.g., an image containing the subject and objects surrounding the subject). The behavior recognition unit (110) recognizes the subject's behavior based on information regarding the objects, information regarding the posture, and information regarding the face, or a combination thereof. Here, the subject's behavior does not simply mean the subject's movements, but also includes the subject's situation. For example, based on an image of the subject sitting at a desk and typing on a keyboard connected to a computer, the behavior recognition unit (110) does not simply recognize the behavior of "the subject sitting," but recognizes the behavior combined with a specific situation such as "the subject working at the desk using a computer." This is possible because the behavior recognition unit (110) recognizes various objects including a subject, such as a desk, chair, monitor, and keyboard, based on an input image, and can recognize the relationship between objects (the monitor and keyboard are placed on the desk) and the relationship between an object and a subject (the subject's hand is placed on the keyboard). Information on the relationship between objects is generated when two related objects coexist or when two objects are within a predetermined distance, and the relationship between an object and a subject can be generated when an object and a subject are within a predetermined distance or are in contact with each other.
[0038] The behavior recognition unit (110) can recognize the behavior of the subject using a behavior recognition model based on information about objects around the subject and information about the subject's posture. For example, based on objects around the subject and the subject's posture, the behavior recognition unit (110) can recognize behaviors such as the subject drinking water from a cup, the subject using a mobile phone, or the subject reading a book while lying down.
[0039] Although not illustrated in FIG. 1, the action recognition unit (110) may include a camera and may generate the input image using the camera. The input image may be an image including a subject and objects around the subject.
[0040] The behavior recognition unit (110) includes a posture detection module (111) and an object detection module (112), and may further include a face keypoint estimation module (113).
[0041] The information regarding the above posture includes the subject's posture and the direction of the subject's movement (body direction). The posture detection module (111) included in the behavior recognition unit (110) recognizes a plurality of key points on the subject's body using a posture estimation model included in the behavior recognition model based on an input image, and determines the subject's posture and the direction of the subject's movement based on the plurality of key points.
[0042] The information regarding the above object includes the recognition result of an object around the subject and the location (coordinates) of the object. The object detection module (112) included in the behavior recognition unit (110) recognizes an object based on an input image using an RCNN (Region-based Convolutional Neural Network)-based object detection model included in the behavior recognition model and determines the location of the recognized object. The object detection model may include a classifier for object recognition.
[0043] The information regarding the face above includes the subject's face direction, gaze, and emotion. The face keypoint estimation module (113) included in the behavior recognition unit (110) generates coordinates of the subject's face features using a face feature recognition model included in the behavior recognition model based on an input image, and determines the subject's face direction, gaze, and emotion based on the coordinates of the face features.
[0044] The information generated by the posture detection module (111), object detection module (112), and face keypoint estimation module (113) included in the action recognition unit (110) is summarized in Table 1. The specific operation of each module (111, 112, 113) included in the action recognition unit (110) will be described later.
[0045] Module Information generated Attitude detection module (111) Information regarding the patient's posture, e.g., patient's posture, direction of movement (body direction), etc. Object detection module (112) Information about objects surrounding the subject (e.g., objects, locations of objects, relationships between objects, relationships between objects and the subject, etc.) Face keypoint estimation module (113) Information regarding the subject's face (e.g., coordinates of facial features, facial direction, gaze, emotion, etc.)
[0046] The behavior recognition unit (110) transmits information regarding the behavior of the recognized subject, as described above, to the service provision unit (120). Additionally, the behavior recognition unit (110) may transmit any one of the information regarding the object, the information regarding the posture, and the information regarding the face, or a combination thereof, to the service provision unit (120).
[0047] Below, the specific operation of each module included in the action recognition unit (110) is described.
[0048] The posture detection module (111) recognizes multiple key points on the subject's body using a posture estimation model included in the behavior recognition model based on an input image, and determines the subject's posture and the direction of the subject's movement based on the multiple key points. For example, the posture estimation model can be implemented using an ML Kit, in which case the posture estimation model recognizes a total of 33 key points (key points) on the subject's body. The posture detection module (111) determines the subject's posture by capturing important joints and body parts such as the head, shoulders, elbows, hips, knees, and feet using the posture estimation model. Additionally, the posture detection module (111) can infer the direction the subject's body (torso) is facing and the direction the subject is moving based on the spatial relationship between the key points. For example, the posture detection module (111) can determine the direction of the subject's torso based on the relative positions of the shoulders and hips, and determine the direction of the subject's movement based on the position of the subject's feet. Identifying the direction the subject is facing is important in the context of human-robot interaction. That is, the result of recognizing whether the subject is facing the robot equipped with the human-robot interaction device (100) or facing the opposite direction can have a significant impact on the decision process, such as whether to provide services to the human-robot interaction device (100). If the subject is looking at the robot, the human-robot interaction device (100) can interpret this as an intention for the subject's potential participation or interaction. On the other hand, if the subject is looking away, the human-robot interaction device (100) can interpret this as indifference or an intention to move in a different direction.
[0049] Furthermore, recognizing the subject's orientation can help the robot predict their behavior and make preemptive decisions. For example, if the subject's posture indicates an intention to rotate toward the robot, the robot can prepare for potential interaction and may slow down or stop if the subject is moving. Ultimately, recognizing key points on the subject's body and accurately interpreting their intentions forms the core of effective human-robot interaction. This ensures that the robot operates in harmony with human intentions and serves as the foundation for more flexible and intuitive interactions.
[0050] The object detection module (112) recognizes an object using an object detection model based on an RCNN (Region-based Convolutional Neural Network) included in an action recognition model based on an input image, and determines the location of the recognized object.
[0051] The above object detection model can be implemented by integrating a model provided by ML Kit that utilizes technology derived from Region-based Convolutional Neural Networks (RCNN). RCNN is a groundbreaking object detection method that combines region proposal and convolutional neural networks. Instead of treating object detection as a single regression problem, RCNN first identifies potential object binding regions in an image using a method called a Region Proposal Network (RPN). The object detection module (112) classifies objects and refines bounding box coordinates using a CNN based on the identified object binding regions. For reference, RPN is a neural network that scans an image and proposes candidate object boundary regions. RPN slides a small window across the entire image and predicts bounding boxes and associated object scores at each location. One of the main strengths of RCNN and its variations is the flexibility of the classifier. That is, by replacing the classifier, the model can be tuned to operate in various environments and detect various object categories. ML Kit's variation of RCNN maintains this versatility, ensuring that the object detection model is adaptable and can operate in a wide range of scenarios.
[0052] The face keypoint estimation module (113) generates coordinates of the subject's face features using a face feature recognition model included in a behavior recognition model based on an input image, and determines the subject's face direction, gaze, and emotion based on the coordinates of the face features.
[0053] The face keypoint estimation module (113) performs an important function in human-robot interaction. The face keypoint estimation module (113) can accurately determine the coordinates of important features such as the subject's eyes, ears, cheeks, nose, and mouth using a face feature recognition model. This detailed recognition helps the robot understand not only the direction of the subject's face but also the location where the subject focuses their gaze. In addition to face feature recognition, the face keypoint estimation module (113) maps the contours of the face and the eyes, eyebrows, lips, and nose in detail. Through this in-depth analysis, the face keypoint estimation module (113) can understand subtle differences in facial expressions and detect and interpret subtle emotional changes. Emotions are essential for interaction between the subject and the robot. The face keypoint estimation module (113) can determine the subject's emotional state by recognizing specific facial expressions, such as a smile or closed eyes. Using the results of this determination, the robot can provide services suitable for the subject's emotions and expectations.
[0054] Consistency is paramount in video streams, and the face keypoint estimation module (113) maintains consistent interaction by tracking faces throughout video frames and assigning a unique identifier to each frame. This continuous tracking ensures that the robot's response and service remain consistent even if the subject changes position or orientation. Additionally, the face keypoint estimation module (113) focuses on real-time video frame processing to ensure smooth and fast operation. This means the robot can immediately interpret facial signals and respond in real-time, reflecting the speed of human interaction and ensuring seamless communication. Consequently, through the advanced capabilities of the face keypoint estimation module (113), the robot can gain a deeper understanding of human users. By keenly interpreting facial cues and emotions, the robot can engage in not only responsive but also empathetic interactions, thereby creating a more intuitive and natural human-robot interaction experience.
[0055] The service providing unit (120) determines the service to be provided to the subject in response to the subject's behavior. The behavior of the subject refers to the behavior recognized by the behavior recognition unit (110).
[0056] The service provider (120) determines whether a service is needed for a subject based on any one of the information regarding the posture, the information regarding the object, and the information regarding the face, or a combination thereof.
[0057] Specifically, the service providing unit (120) determines whether there is an intention to request a service based on any one of the information regarding the posture and the information regarding the face, or a combination thereof. For example, the service providing unit (120) can determine whether the subject has an intention to request a service based on any one of the subject's direction of movement (direction of the body), the subject's face direction and gaze, or a combination thereof.
[0058] If the service providing unit (120) determines that the subject has an intention to request a service, it determines the service to be provided to the subject in response to the subject's behavior. If the service providing unit (120) determines that the subject does not have an intention to request a service, it determines whether a service is needed for the subject by using a pre-set table or a pre-trained neural network based on the subject's behavior. For example, if the subject maintains a lying position for a long time without covering with a blanket or using a pillow (possibility of an emergency situation), the service providing unit (120) may determine that a service is needed for the subject.
[0059] The service providing unit (120) determines the service to be proposed to the subject in response to the subject's behavior. The service providing unit (120) can determine the service to be proposed to the subject by inputting the subject's behavior (e.g., embedding vector) into a pre-trained neural network (behavior-service model). The service providing unit (120) may provide the service immediately without going through the process of proposing the service to the subject, depending on the subject's behavior or the determined service. For example, if the subject's behavior corresponds to an emergency situation or the determined service is 'sending a text message requesting help,' the service providing unit (120) provides the determined service without going through the process of proposing the service to the subject.
[0060] Since the behavior recognition unit (110) specifically recognizes the behavior of the subject, the service provision unit (120) can determine an appropriate service based on the subject's behavior. For example, even if the subject is in a lying position, the behavior recognition unit (110) combines information about the object and information about the posture to distinguish between the behavior of the subject reading a book while lying down and the behavior of the subject watching TV while lying down. Therefore, the service provision unit (120) can provide a service that plays quiet classical music when the subject is reading a book while lying down, and a service that brings water or snacks to the subject when the subject is watching TV while lying down, and can provide a service suitable for the subject's behavior.
[0061] The speech unit (130) proposes a service to the subject through voice guidance. The speech unit (130) generates and speaks a guidance voice that proposes a service to the subject, as determined by the service provider (120), through voice synthesis. For example, if the subject is reading, the speech unit (130) can speak a guidance voice saying, "It looks like you are reading. Shall I play some music suitable for reading?"
[0062] The voice recognition unit (140) collects the subject's voice for a predetermined period of time after the guidance voice utterance of the speech unit (130), and determines whether the subject's voice corresponds to a positive, negative, or neutral response through voice recognition and analysis. That is, the voice recognition unit (140) determines the subject's response based on the subject's voice. If the subject's voice is absent (no response) or if it is difficult to determine whether it is positive or negative (impossible to determine), the voice recognition unit (140) classifies the subject's response as neutral.
[0063] When the voice recognition unit (140) determines the subject's response, the service provision unit (120) determines whether to provide the service based on the subject's response. If the subject's response is positive, the service provision unit (120) determines that the service is provided. Additionally, if the subject's response is negative, the service provision unit (120) determines that the service is not provided. Additionally, if the subject's response is neutral, the service provision unit (120) determines whether to provide the service based on the type of proposed service or the subject's behavior. For example, if the proposed service is related to 'daily health' (e.g., brushing teeth), the service provision unit (120) decides to provide a health guidance service (e.g., "You must brush your teeth for more than 3 minutes to help your dental health") even if the subject's response is neutral. In another example, the service provider (120) decides to provide a service that counts the number of push-ups performed by the subject and notifies the subject, even if the subject's response is neutral, when the subject's behavior corresponds to 'exercise' (e.g., push-ups).
[0064] When the service provider (120) decides to provide a service to the subject, it provides the service using an external device that is mounted on the human-robot interaction device (100) or connected via a network. For example, the service provider (120) may provide guidance on exercise or health through the speech unit (130), and may provide music or video through a mobile device that includes the human-robot interaction device (100) or is connected to the human-robot interaction device (100). Additionally, if the subject's behavior corresponds to an emergency situation, the service provider (120) may send a text message requesting help to a registered guardian through the connected mobile device.
[0065] The model training unit (150) trains a neural network-based model (e.g., behavior recognition model, behavior-service model) used by the human-robot interaction device (100) based on an externally input training dataset and delivers it to each component of the human-robot interaction device (100). For example, the model training unit (150) trains a behavior recognition model and delivers it to the behavior recognition unit (110). Additionally, the model training unit (150) can deliver the quantized model to each component of the human-robot interaction device (100) after quantizing the neural network-based model by applying quantization techniques, including Quantization-Aware Training (QAT), so that the neural network-based model can operate efficiently on a mobile device. In neural networks, quantization generally refers to reducing memory and computation costs by reducing the precision of weights and activation functions. Mobile devices with limited computational resources and memory may struggle to execute large-scale deep learning models. Furthermore, power consumption associated with the execution of such models can be a significant issue for battery-operated mobile devices. While reducing the size of the neural network is one solution, it often results in a decrease in model accuracy; however, quantization provides a balance between the size of the neural network and model accuracy. That is, quantization reduces the model size and computational requirements without significantly degrading model performance. QAT (Quantization Aware Training) is a technique that ensures the quantized model maintains the accuracy of the original model. Additionally, the model training unit (150) can efficiently deploy deep learning models (neural network models) on mobile platforms using TensorFlow Lite.
[0066] FIG. 2 is a flowchart illustrating a human-robot interaction method according to an embodiment of the present invention. For convenience of explanation, it is assumed that the human-robot interaction method according to an embodiment of the present invention is performed by a human-robot interaction device (100).
[0067] Referring to FIG. 2, a human-robot interaction method according to one embodiment of the present invention comprises steps S210 through S270. The human-robot interaction method illustrated in FIG. 2 is according to one embodiment, and the steps of the human-robot interaction method according to the present invention are not limited to the embodiment illustrated in FIG. 2 and may be added, changed, or deleted as necessary. For example, the human-robot interaction method according to one embodiment of the present invention may be configured to include only steps S210 and S230.
[0068] Step S210 is the step of recognizing the subject's behavior.
[0069] The behavior recognition unit (110) generates information regarding one or more objects around the subject, information regarding the subject's posture, and information regarding the subject's face, or a combination thereof, using a neural network-based behavior recognition model based on an input image (e.g., an image containing the subject and objects surrounding the subject). The behavior recognition unit (110) recognizes the subject's behavior based on information regarding the objects, information regarding the posture, and information regarding the face, or a combination thereof.
[0070] Step S220 is the step of determining whether the subject requires service.
[0071] The service provider (120) determines whether a service is needed for a subject based on any one of the information regarding the posture, the information regarding the object, and the information regarding the face, or a combination thereof.
[0072] Specifically, the service providing unit (120) can determine whether there is an intention to request a service based on any one of the information regarding the posture and the information regarding the face, or a combination thereof. For example, the service providing unit (120) can determine whether the subject has an intention to request a service based on any one of the subject's direction of movement (direction of the body), the subject's face direction and gaze, or a combination thereof.
[0073] If the service provider (120) determines that the subject has an intention to request a service, it proceeds to step S230. If the service provider (120) determines that the subject does not have an intention to request a service, it determines whether a service is needed for the subject by using a pre-set table or a pre-trained neural network based on the subject's behavior. For example, if the subject maintains a lying position for a long time without covering with a blanket or using a pillow (possibility of an emergency situation), the service provider (120) may determine that a service is needed for the subject.
[0074] The service provider (120) terminates the process if it determines that the service for the subject is not needed.
[0075] Step S230 is the step of determining the service to propose to the target.
[0076] The service provider (120) determines the service to be proposed to the subject in response to the subject's behavior. The service provider (120) can determine the service to be proposed to the subject by inputting the subject's behavior (e.g., embedding vector) into a pre-trained neural network (behavior-service model). Depending on the subject's behavior or the determined service, the service provider (120) may proceed directly to step S270 without going through steps S240 through S260. For example, if the subject's behavior corresponds to an emergency situation or the determined service is 'sending a text message requesting help', the service provider (120) proceeds to step S270.
[0077] Step S240 is the step of proposing a service to the target person.
[0078] The speech unit (130) generates and speaks a guidance voice suggesting a service determined by the service provider (120) to the subject through speech synthesis. For example, if the subject is reading, the speech unit (130) can speak a guidance voice saying, "It looks like you are reading. Shall I play some music suitable for reading?"
[0079] Step S250 is the step of collecting the subject's response.
[0080] The voice recognition unit (140) collects the subject's voice for a predetermined period of time after the guidance voice utterance of the speech unit (130), and determines whether the subject's voice corresponds to a positive, negative, or neutral response through voice recognition and analysis. That is, the voice recognition unit (140) determines the subject's response based on the subject's voice. If the subject's voice is absent (no response) or if it is difficult to determine whether it is positive or negative (impossible to determine), the voice recognition unit (140) classifies the subject's response as neutral.
[0081] Step S260 is the step for determining whether to provide the service.
[0082] The service provider (120) determines whether to provide the service based on the subject's response. If the subject's response is positive, the service provider (120) determines that the service is provided. Additionally, if the subject's response is negative, the service provider (120) determines that the service is not provided. Additionally, if the subject's response is neutral, the service provider (120) determines whether to provide the service based on the type of proposed service or the subject's behavior. For example, if the proposed service is related to 'daily health' (e.g., brushing teeth), the service provider (120) decides to provide a health guidance service (e.g., "You must brush your teeth for at least 3 minutes to help your dental health") even if the subject's response is neutral. As another example, if the subject's behavior is related to 'exercise' (e.g., push-ups), the service provider (120) decides to provide a service that informs the subject of the results of the exercise performed (e.g., number of push-ups) even if the subject's response is neutral.
[0083] The service provider (120) proceeds to step S270 if it determines that it is providing the result service of step S260, and terminates the process otherwise.
[0084] Step S270 is the step of providing services.
[0085] The service provider (120) provides the service determined in step S230. The service provider (120) provides the service using an external device that is mounted on the human-robot interaction device (100) or connected via a network. For example, the service provider (120) may provide guidance on exercise or health through the speech unit (130), and may provide music or video through a mobile device that includes the human-robot interaction device (100) or is connected to the human-robot interaction device (100). Additionally, if the subject's behavior corresponds to an emergency situation, the service provider (120) may send a text message requesting help to a registered guardian through the connected mobile device.
[0086] The aforementioned human-robot interaction method has been described with reference to the flowchart presented in the drawings. For simplicity of explanation, the method has been illustrated and described in a series of blocks; however, the present invention is not limited to the order of said blocks, and some blocks may occur in a different order or simultaneously with other blocks as illustrated and described herein, and various other branches, flow paths, and sequences of blocks may be implemented to achieve the same or similar results. Furthermore, not all illustrated blocks may be required for the implementation of the method described herein.
[0087] Meanwhile, in the description with reference to FIG. 2, each step may be further divided into additional steps or combined into fewer steps according to an embodiment of the present invention. Also, some steps may be omitted as necessary, and the order between steps may be changed. Furthermore, even if other omitted details are included, the contents of FIG. 1 may be applied to the contents of FIG. 2. Also, the contents of FIG. 2 may be applied to the contents of FIG. 1.
[0088] FIG. 3 is a screen showing an experiment on behavior recognition of a human-robot interaction device and method according to an embodiment of the present invention. FIG. 3 is a screenshot of a screen showing the operation of a human-robot interaction device implemented according to an embodiment of the present invention on a Galaxy A24. At the time of the experiment, the smartphone was a representative product of mainstream and mid-range smartphones. While the smartphone has practical specifications, it does not possess the computing power of a high-end smartphone. The experiment was conducted for the purpose of verifying whether the human-robot interaction device and method according to the present invention could operate efficiently on the smartphone. The device (smartphone) incorporating the configuration of the present invention achieved a frame rate of 40 FPS, which is near real-time, in most human-robot interaction scenarios. The image in FIG. 3 shows a subject sitting in a chair and using a computer, and the device successfully detected various objects in the field, such as the subject, the desk, the computer, and the keyboard. Furthermore, the device clearly recognized the subject's behavior (working at the desk using the computer) based on information regarding objects around the subject and information regarding the subject's posture.
[0089] FIG. 4 is a block diagram showing a computer system for implementing a human-robot interaction method according to an embodiment of the present invention. A human-robot interaction device (100) according to an embodiment of the present invention may be implemented in the form of FIG. 4.
[0090] Referring to FIG. 4, a computer system (1000) may include at least one of a processor (1010), memory (1030), input interface device (1050), output interface device (1060), and storage device (1040) communicating via a bus (1070). The computer system (1000) may also further include a communication device (1020) coupled to a network. The processor (1010) may be a central processing unit (CPU) or a semiconductor device that executes instructions stored in memory (1030) or storage device (1040). Memory (1030) and storage device (1040) may include various forms of volatile or non-volatile storage media. For example, memory may include read-only memory (ROM) and random access memory (RAM). In the embodiments of this description, memory may be located inside or outside the processor, and memory may be connected to the processor through various known means. Memory is a volatile or non-volatile storage medium of various forms, and for example, memory may include read-only memory (ROM) or random access memory (RAM).
[0091] Accordingly, embodiments of the present invention may be implemented as a method implemented on a computer or as a non-transient computer-readable medium storing computer-executable instructions. In one embodiment, when executed by a processor, the computer-readable instructions may perform a method according to at least one aspect of the present description.
[0092] The communication device (1020) can transmit or receive wired or wireless signals.
[0093] In addition, the method according to an embodiment of the present invention may be implemented in the form of program instructions that can be executed through various computer means and may be recorded on a computer-readable medium.
[0094] The above computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the computer-readable medium may be specially designed and configured for embodiments of the present invention, or they may be known and available to a person skilled in the art of computer software. The computer-readable recording medium may include a hardware device configured to store and execute program instructions. For example, the computer-readable recording medium may be magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; ROM; RAM; flash memory, etc. The program instructions may include not only machine code, such as that generated by a compiler, but also high-level language code that can be executed by a computer through an interpreter, etc.
[0095] For reference, the components according to the embodiments of the present invention may be implemented in the form of software or hardware such as a DSP (digital signal processor), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit), and may perform certain roles.
[0096] However, 'components' are not limited to software or hardware, and each component may be configured to reside in an addressable storage medium or configured to operate one or more processors.
[0097] Accordingly, as an example, components include components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.
[0098] Components and the functions provided within them can be combined into a smaller number of components or further separated into additional components.
[0099] Meanwhile, it will be understood that each block of the flowcharts and combinations of the flowcharts can be executed by computer program instructions. Since these computer program instructions can be loaded onto the processor of a general-purpose computer, a computer for special purposes, or other programmable data processing equipment, the instructions executed through the processor of the computer or other programmable data processing equipment create a means to perform the functions described in the flowchart block(s). Since computer program instructions can also be loaded onto a computer or other programmable data processing equipment, the instructions that execute a series of operation steps on the computer or other programmable data processing equipment to create a process executed by the computer can also provide steps for executing the functions described in the flowchart block(s).
[0100] Additionally, each block may represent a module, segment, or part of code containing one or more executable instructions for executing a specified logical function(s). It should also be noted that in some alternative execution examples, the functions mentioned in the blocks may occur out of order. For instance, two blocks described in succession may actually be executed substantially simultaneously, or the blocks may be executed in reverse order according to their corresponding functions.
[0101] The terms 'part' or 'module' used in this embodiment refer to software or hardware components such as FPGAs or ASICs, and the 'part' or 'module' performs certain roles. However, the meaning of 'part' or 'module' is not limited to software or hardware. The 'part' or 'module' may be configured to reside in an addressable storage medium or may be configured to run one or more processors. Accordingly, as an example, the 'part' or 'module' includes components such as software components, object-oriented software components, class components, and task components, as well as processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables. The functions provided within the components and 'parts' or 'modules' may be combined into a smaller number of components and 'parts' or 'modules', or further separated into additional components and 'parts' or 'modules'. In addition, the components and 'parts' or 'modules' may be implemented to utilize one or more CPUs within the device or secure multimedia card.
[0102] Although the present invention has been described above with reference to preferred embodiments, those skilled in the art will understand that various modifications and changes can be made to the invention without departing from the spirit and scope of the invention as described in the following claims. Explanation of the symbols
[0103] 100: Human-Robot Interaction Device 110: Behavior recognition unit 111: Attitude detection module 112: Object detection module 113: Face Keypoint Estimation Module 120: Service Provider 130: Ignition part 140: Voice recognition unit 150: Model Training Department 1000: Computer System 1010: Processor 1020: Communication device 1030: Memory 1040: Storage device 1050: Input interface device 1060: Output interface device 1070: Bus
Claims
Claim 1 An action recognition unit that, based on an input image, uses a neural network-based action recognition model to generate information regarding one or more objects surrounding a subject, information regarding the subject's posture, and information regarding the subject's face, and recognizes the subject's behavior associated with the subject's situation based on the information regarding the objects, the information regarding the posture, and the information regarding the face; and a model training unit that generates the action recognition model by training a neural network-based model based on an input training dataset, applies Quantization-Aware Training (QAT) to the action recognition model to quantize it, and then transmits it to the action recognition unit.and includes a service providing unit that determines a service to be provided to the subject in response to the behavior of the subject, wherein information regarding the object includes the recognition result of the object, the location of the object, relationship information between the objects, and relationship information between the object and the subject, wherein information regarding the posture includes the posture of the subject and the direction of the subject's movement, and information regarding the face includes the direction of the subject's face, gaze, and emotion, and the behavior recognition unit recognizes the object and determines the location of the object using an RCNN-based object detection model included in the behavior recognition model based on the input image, and if there are objects within a predetermined distance among the objects, generates relationship information between the objects indicating the positional relationship of the objects, and if the object and the subject are within a predetermined distance or are in contact with each other, generates relationship information between the object and the subject indicating the positional relationship between the object and the subject, and recognizes a plurality of key points on the body of the subject using a posture estimation model included in the behavior recognition model based on the input image, determines the posture of the subject and the direction of the subject's movement based on the plurality of key points, and the face included in the behavior recognition model based on the input image A human-robot interaction device using a neural network, wherein the device generates coordinates of the subject's facial features using a feature recognition model and determines the subject's facial direction, gaze, and emotion based on the coordinates of the facial features; the service providing unit determines whether a service is needed for the subject based on information regarding the posture, and if it is determined that a service is needed for the subject, determines a service to be provided to the subject in response to the subject's behavior and emotion. Claim 2 delete Claim 3 delete Claim 4 delete Claim 5 delete Claim 6 delete Claim 7 A method performed by a human-robot interaction device comprises: a step of generating an action recognition model by training a neural network-based model based on an input learning dataset and quantizing the action recognition model by applying Quantization-Aware Training (QAT); a step of inputting an input image into the action recognition model to generate information regarding one or more objects surrounding a subject, information regarding the subject's posture, and information regarding the subject's face, and recognizing the subject's behavior based on the information regarding the objects, the information regarding the posture, and the information regarding the face; and a step of determining whether a service is required for the subject based on the information regarding the posture. The method includes a step of determining a service to be provided to the subject in response to the subject's behavior when it is determined that a service is required for the subject, wherein the information regarding the object includes the recognition result of the object, the location of the object, relationship information between the objects, and relationship information between the object and the subject, wherein the information regarding the posture includes the posture of the subject and the direction of the subject's movement, and the information regarding the face includes the direction of the subject's face, gaze, and emotion, and the step of recognizing the behavior includes: a step of recognizing the object and determining the location of the object using an RCNN-based object detection model included in the behavior recognition model based on the input image; a step of generating relationship information between the objects indicating the positional relationship of the objects when there are objects within a predetermined distance among the objects; and a step of generating relationship information between the object and the subject indicating the positional relationship between the object and the subject when the object and the subject are within a predetermined distance or are in contact with each other.A human-robot interaction method using a neural network, comprising: a step of recognizing a plurality of key points on the body of a subject using a pose estimation model included in the behavior recognition model based on the input image, and determining the subject's pose and the direction of the subject's movement based on the plurality of key points; and a step of generating coordinates of the subject's facial features using a face feature recognition model included in the behavior recognition model based on the input image, and determining the subject's facial direction, gaze, and emotion based on the coordinates of the facial features, wherein the step of determining the service includes determining the service to be provided to the subject in response to the subject's behavior and emotion. Claim 8 delete