Methods, devices, media, and procedures for waking up robots.
By fusing dynamic and static gesture recognition methods with affine transformations and lightweight convolutional neural network models on the upper body image frames of a humanoid robot, the problem of gesture wake-up on curved screens being easily affected by lighting conditions is solved, thereby improving the wake-up success rate and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI MATRIX SUPER INTELLIGENT SYSTEM INTEGRATION CO LTD
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-26
AI Technical Summary
In humanoid robot scenarios based on dark curved screen perception, gesture wake-up is easily affected by the lighting environment, resulting in a high false wake-up rate and affecting the user interaction experience.
An improved dynamic gesture recognition method is adopted, which combines dynamic and static gesture recognition by performing affine transformation on a preset number of upper body image frames and fusing them with a lightweight convolutional neural network model. This reduces the false wake-up rate.
It significantly reduced the false wake-up rate of gestures and improved the user interaction experience, especially in low light, bright light and sports modes, and improved the success rate of gesture wake-up.
Smart Images

Figure CN121468570B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics, and more particularly to a technique for waking up a robot. Background Technology
[0002] In existing technologies, humanoid robots can be woken up and enter task mode simply by a user's natural gesture in low-power standby mode, eliminating the need for physical buttons, voice commands, or remote controls. Traditional buttons / voice commands require proximity to the device or the use of sound, resulting in a poor user experience in situations where hands are occupied, noisy environments, or social settings. Gestures, on the other hand, offer the advantages of being "non-contact, intuitive, and silent," and are considered the next generation of natural human-machine interfaces. Humanoid robots are often deployed in homes, showrooms, and collaborative production lines, and users expect to wake them up "like greeting a person." Gesture wake-up has become a key function for a "human-like" experience. Since robots need to remain in standby mode for extended periods but cannot operate all sensors at full power, gesture wake-up modules typically operate with "ultra-low power sensing + lightweight detection," only triggering the main system to power on after confirming intent, thus balancing battery life and response speed.
[0003] In existing technologies, for the sake of anthropomorphism and a sense of technology, dark curved screens are usually used to cover the cameras of humanoid robots. In humanoid robot scenarios based on curved screen perception, the camera is placed inside the semi-transparent curved screen. External light enters the camera sensor after two refractions / reflections, resulting in a 20%-40% decrease in brightness. Color shift and backlighting are dominant, causing the images captured by the humanoid robot's camera to be darker than when it is not wearing a curved screen. At the same time, the curvature of the curved screen introduces radial distortion and local reflections. Under these circumstances, existing technologies are prone to accidental wake-up by gestures, significantly increasing the wake-up time, and are highly susceptible to the influence of the lighting environment, affecting the user interaction experience. Summary of the Invention
[0004] One object of this application is to provide a method, apparatus, medium, and program product for waking up a robot.
[0005] According to one aspect of this application, a method for waking up a robot is provided, the method comprising:
[0006] Based on the video stream obtained by the robot, obtain the coordinates of the upper body human body rectangle corresponding to at least one human body in each video frame of the video stream;
[0007] Affine transformation is performed on each video frame based on the coordinates of the upper body human body rectangle to obtain an upper body image frame of each video frame with respect to the at least one human body.
[0008] Dynamic gesture recognition is performed on the upper body image frames to obtain the corresponding dynamic gesture recognition results. If the cumulative number of upper body image frames reaches the first preset number of frames and the dynamic gesture recognition result corresponding to the current upper body image frame is wake-up, static gesture detection is performed on the currently accumulated multiple upper body image frames. If the number of upper body image frames in the multiple upper body image frames that have been detected with preset gestures meets the preset conditions, the wake-up operation is performed on the robot.
[0009] According to one aspect of this application, a computer device for waking up a robot is provided, comprising a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of any of the methods described above.
[0010] According to one aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of any of the methods described above.
[0011] According to one aspect of this application, a computer program product is provided, comprising a computer program, characterized in that, when executed by a processor, the computer program implements the steps of any of the methods described above.
[0012] According to one aspect of this application, a computer device for waking up a robot is provided, the device comprising:
[0013] The module is used to obtain the coordinates of the upper body human body rectangle corresponding to at least one human body in each video frame of the video stream obtained by the robot.
[0014] The first and second modules are used to perform an affine transformation on each video frame based on the coordinates of the upper body human body rectangle to obtain an upper body image frame of each video frame with respect to the at least one human body.
[0015] The first and third modules are used to perform dynamic gesture recognition on the upper body image frames to obtain the corresponding dynamic gesture recognition results. If the cumulative number of upper body image frames reaches the first preset number of frames and the dynamic gesture recognition result corresponding to the current upper body image frame is wake-up, static gesture detection is performed on the currently accumulated multiple upper body image frames. If the number of upper body image frames in the multiple upper body image frames that have been detected with preset gestures meets the preset conditions, the robot is woken up.
[0016] Compared with existing technologies, this application proposes an improved dynamic gesture recognition method suitable for edge device operation for humanoid robots based on dark curved screen perception. Without increasing parameters and computational load, it fuses the features of upper body image frames of a preset number of frames, i.e., simultaneously modeling space and time. This significantly reduces the occurrence rate of gesture false wake-up in humanoid robot scenarios with curved screen perception, improving the user interaction experience. It also improves the gesture wake-up success rate of humanoid robots in motion mode and significantly reduces the gesture false wake-up rate. Furthermore, in low light, bright light, and motion modes, since lighting or motion will make the image darker and blurrier, this application adds static gesture detection to further reduce gesture false wake-up, which can improve the gesture wake-up success rate of humanoid robots in low light and bright light environments and significantly reduce the gesture false wake-up rate. Attached Figure Description
[0017] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0018] Figure 1 This diagram illustrates a method for waking up a robot according to one embodiment of the present application.
[0019] Figure 2 This diagram illustrates a method for dynamic gesture recognition according to an embodiment of the present application.
[0020] Figure 3 This diagram illustrates a method for static gesture recognition according to an embodiment of the present application.
[0021] Figure 4 This diagram illustrates a computer device structure for waking up a robot according to one embodiment of the present application.
[0022] Figure 5 Exemplary systems that can be used to implement the various embodiments described in this application are shown.
[0023] The same or similar reference numerals in the accompanying drawings represent the same or similar parts. Detailed Implementation
[0024] The present application will now be described in further detail with reference to the accompanying drawings.
[0025] In a typical configuration of this application, the terminal, the device of the service network, and the trusted party all include one or more processors (e.g., a central processing unit (CPU)), input / output interfaces, network interfaces, and memory.
[0026] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory. Memory is an example of computer-readable media.
[0027] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PCM), programmable random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0028] The devices referred to in this application include, but are not limited to, user equipment, network equipment, or devices composed of user equipment and network equipment integrated through a network. The user equipment includes, but is not limited to, any mobile electronic product capable of human-computer interaction (e.g., via a touchpad), such as smartphones and tablets. These mobile electronic products can use any operating system, such as Android or iOS. The network equipment includes an electronic device capable of automatically performing numerical calculations and information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), and embedded devices. The network equipment includes, but is not limited to, computers, network hosts, single network servers, multiple network server clusters, or a cloud composed of multiple servers. Here, the cloud consists of a large number of computers or network servers based on cloud computing, where cloud computing is a type of distributed computing, consisting of a virtual supercomputer composed of a group of loosely coupled computer clusters. The network includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, VPN network, wireless ad hoc network, etc. Preferably, the device can also be a program running on the user equipment, network device, or a device formed by integrating user equipment and network device, network device, touch terminal, or network device and touch terminal through a network.
[0029] Of course, those skilled in the art should understand that the above-described devices are merely examples, and other existing or future devices that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0030] In the description of this application, "multiple" means two or more, unless otherwise expressly and specifically defined.
[0031] Figure 1The diagram illustrates a method for waking up a robot according to an embodiment of this application. The method includes steps S11, S12, and S13. In step S11, a computer device obtains the upper body bounding box coordinates of at least one human body in each video frame of the video stream obtained by the robot. In step S12, the computer device performs an affine transformation on each video frame based on the upper body bounding box coordinates to obtain an upper body image frame of each video frame with respect to the at least one human body. In step S13, the computer device performs dynamic gesture recognition on the upper body image frames to obtain corresponding dynamic gesture recognition results. If the cumulative number of upper body image frames reaches a first preset frame number, and the dynamic gesture recognition result corresponding to the current upper body image frame is "wake-up," static gesture detection is performed on the currently accumulated multiple upper body image frames. If the number of upper body image frames in the multiple upper body image frames that have detected a preset gesture meets a preset condition, a wake-up operation is performed on the robot.
[0032] In step S11, the computer device obtains the coordinates of the upper body rectangle corresponding to at least one human body in each video frame of the video stream obtained by the robot.
[0033] In some embodiments, the computer device may be the robot itself, or it may be a remote workstation responsible for model reasoning.
[0034] In some embodiments, the robot performs target (here, the target only includes human bodies) tracking on the video stream obtained by its camera, obtaining the identification information (e.g., ID) of at least one human body contained in each video frame of the video stream, and the coordinates of the bounding box corresponding to the upper body of each human body in each video frame, i.e., the coordinates of the upper body human body bounding box. In some embodiments, the upper body human body bounding box coordinates include the coordinates of the four vertices of the upper body human body bounding box, or, only the coordinates of the two diagonal vertices of the upper body human body bounding box, for example, only the coordinates of the upper left corner and the lower right corner, or only the coordinates of the lower left corner and the upper right corner.
[0035] In step S12, the computer device performs an affine transformation on each video frame based on the coordinates of the upper body human body rectangle to obtain an upper body image frame of each video frame with respect to the at least one human body.
[0036] In some embodiments, an affine transformation is a linear geometric transformation that preserves the parallel lines and the proportions of collinear points. It combines scaling, rotation, cropping, and translation effects by applying linear transformations and translation operations to the vertices of an image or graphic. In some embodiments, for each video frame in a video stream, an affine transformation (e.g., bilinear interpolation affine transformation) is performed on the video frame based on the coordinates of the upper body bounding box corresponding to each human figure in the video frame. This extracts the image of the ROI (Region of Interest) for each human figure in the video frame, i.e., the upper body image frame. By obtaining the upper body image frame of each human figure in each video frame through the affine transformation, the image area containing the upper body bounding box of each human figure in the video frame is cropped and scaled to a fixed size while maintaining the proportions of the human figures (avoiding image distortion and keeping the target in the original image from deforming).
[0037] In step S13, the computer device performs dynamic gesture recognition on the upper body image frames to obtain the corresponding dynamic gesture recognition result. If the number of upper body image frames reaches the first preset number, and the dynamic gesture recognition result corresponding to the current upper body image frame is wake-up, static gesture detection is performed on the currently accumulated multiple upper body image frames. If the number of upper body image frames in the multiple upper body image frames that have been detected with preset gestures meets the preset condition, the robot is woken up.
[0038] In some embodiments, for each human body contained in the video stream, dynamic gesture recognition is performed on the upper body image frame corresponding to that human body in each video frame to obtain the dynamic gesture recognition result corresponding to that upper body image frame. Here, dynamic gesture recognition is a technology that captures, analyzes, and semantically interprets the motion trajectory, action sequence, and posture change process of a user's hand in a continuous temporal space. The core is to recognize gesture actions with a time dimension, rather than a single fixed posture. It recognizes the continuous motion trajectory and temporal change process of the gesture, which requires analyzing the position and posture evolution law of the hand in multiple consecutive images. Essentially, it is the analysis of a temporal video sequence, that is, the analysis of the upper body image frame sequence composed of the upper body image frames corresponding to each video frame in the video stream.
[0039] In some embodiments, for each human body included in the video stream, the dynamic gesture recognition result corresponding to each upper body image frame in the upper body image frame sequence corresponding to that human body in each video frame can be obtained by inputting the upper body image frame sequence (e.g., MobileNetV2 model, a lightweight convolutional neural network proposed by Google, designed for mobile and embedded devices, and an upgrade of MobileNetV1) into a lightweight convolutional neural network model. The lightweight convolutional neural network model is a type of convolutional neural network (CNN) designed for computationally limited scenarios (such as mobile devices, embedded devices, and edge computing terminals). Its core objective is to achieve fast inference and low resource consumption by simplifying the network structure, reducing the number of parameters and computational load while maintaining a certain level of accuracy. In some embodiments, the dynamic gesture recognition result includes both awake and non-awakened results.
[0040] In some embodiments, for each human body included in the video stream, if the cumulative number of upper body image frames corresponding to that human body reaches a first preset number of frames (e.g., 16 frames), and the dynamic gesture recognition result of the current upper body image frame (e.g., the 16th frame) corresponding to that human body is "wake-up," then static gesture detection is performed on each of the current cumulative upper body image frames (i.e., the 16 upper body image frames). For each upper body image frame, if a preset gesture is recognized from that upper body image frame, the static gesture recognition result corresponding to that upper body image frame is output. The static gesture recognition result includes the coordinates of the rectangular detection box corresponding to the preset gesture in the upper body image frame and the category probability corresponding to the preset gesture. Based on the static gesture recognition result, it can be determined whether the preset gesture has been detected in the upper body image frame. Static gesture detection is a recognition technology for gestures in a static state. It focuses on the spatial morphological features of the gesture, rather than the temporal change features of the action. The object of static gesture detection is the gesture posture at a certain moment, such as an open palm, a clenched fist, or making an "OK" sign. Fixed gestures, such as raising the index finger, can be detected and recognized using only a single image frame, without requiring a continuous sequence of video frames. The core basis for recognition is the spatial geometric features of the gesture, including its shape and outline, the number of fingers, joint positions, and palm orientation. In some embodiments, the yolov5n model (an extremely lightweight version of the YOLOv5 object detection algorithm family developed by the Ultralytics team, designed specifically for edge devices, mobile devices, and low-power hardware with limited computing power; it is the model with the fewest parameters and the fastest inference speed in the YOLOv5 series) can be used to perform static gesture detection on upper body image frames. That is, the upper body image frame is input into the yolov5n model, and the static gesture recognition result corresponding to the upper body image frame is obtained from the network's output. In some embodiments, the rule module judges the static gesture recognition results corresponding to the currently accumulated multiple upper body image frames. If the number of upper body image frames (e.g., 16 upper body image frames) that detect a preset gesture meets a preset condition, it is determined that the robot needs to be woken up and the corresponding wake-up operation is performed on the robot. Otherwise, it is determined that the robot does not need to be woken up. The preset condition includes, but is not limited to, the number of upper body image frames that detect a preset gesture being greater than or equal to a second preset number of frames (e.g., 8 frames), and the ratio of the number of upper body image frames that detect a preset gesture to the total number of frames of the currently accumulated multiple upper body image frames being greater than or equal to a preset ratio threshold. This example embodiment does not impose any special limitations on these conditions. In some embodiments, the rule module can expand and update the rules according to the usage scenario. In some embodiments, at a frame rate of 30fps, 16 frames are approximately equal to 0.5-0.7 seconds of video, which is sufficient to cover a complete cycle of most waving actions, ensuring that the model can capture the start and end of the action, and achieving a trade-off between the temporal receptive field and computational efficiency.In some embodiments, performing a wake-up operation on a robot refers to the process of activating the robot from a dormant or low-power state, enabling it to enter an interactive working state.
[0041] This application proposes an improved dynamic gesture recognition method suitable for edge device operation for humanoid robots based on dark curved screen perception. Without increasing parameters or computational load, it fuses features from a preset number of upper-body image frames, simultaneously modeling space and time. This significantly reduces the incidence of false wake-up gestures in humanoid robot scenarios with curved screen perception, improving the user experience. It also increases the success rate of gesture wake-up in motion modes and significantly reduces the false wake-up gesture rate. Furthermore, in low light, bright light, and motion modes, where lighting or movement makes images darker and blurrier, this application adds static gesture detection to further reduce false wake-up gestures. This improves the success rate of gesture wake-up in low light and bright light environments and significantly reduces the false wake-up gesture rate.
[0042] In some embodiments, obtaining the upper body bounding box coordinates of at least one human body in each video frame of the video stream obtained by the robot includes: performing pedestrian tracking on the video stream obtained by the robot, obtaining the human bounding box coordinates and their corresponding first score for at least one human body in each video frame, and the coordinates of multiple key points and their corresponding second score for the at least one human body; and determining the upper body bounding box coordinates of at least one human body in each video frame based on the human bounding box coordinates, the first score, the key point coordinates, and the second score. In some embodiments, pedestrian tracking is a technique in the field of computer vision that locates and tracks one or more pedestrian targets in a continuous video frame sequence. The core is to establish the correspondence between target pedestrians in different frames to achieve continuous monitoring of their position and movement trajectory. In some embodiments, a pedestrian tracking module based on Ultralytics YOLO (an open-source, high-performance object detection and computer vision algorithm framework maintained by Ultralytics) performs pedestrian tracking on the video stream obtained by the robot's camera. This obtains the identification information (e.g., ID) of at least one human body contained in each video frame of the video stream, the coordinates of the bounding box corresponding to each human body in each video frame (i.e., the coordinates of the human body bounding box), a first score corresponding to the human body bounding box coordinates (i.e., the confidence (probability) that the bounding box contains a real human body), multiple (e.g., 17) keypoint coordinates corresponding to each human body in each video frame, and a second score corresponding to the keypoint coordinates (i.e., the confidence (probability) of the localization accuracy). In some embodiments, based on the human body bounding box coordinates, the first score corresponding to the human body bounding box coordinates, the multiple keypoint coordinates, and the second score corresponding to each keypoint coordinate, the upper body bounding box coordinates of each human body contained in the video stream in each video frame are calculated using coordinates.
[0043] In some embodiments, determining the upper body rectangle coordinates of at least one human body in each video frame based on the human body rectangle coordinates, the first score, the keypoint coordinates, and the second score includes: determining the lower right ordinate of the upper body rectangle of at least one human body in each video frame based on the lower right ordinate of the human body rectangle, the first score, the keypoint coordinates, and the second score, wherein the upper left and lower right horizontal coordinates of the upper body rectangle remain unchanged compared to the human body rectangle coordinates. In some embodiments, the lower right ordinate of the upper body rectangle of each human body in each video frame can be calculated based on the lower right ordinate of the human body rectangle, the first score corresponding to the human body rectangle coordinates, multiple keypoint coordinates, and the second score corresponding to each keypoint coordinate, wherein the upper left (including upper left horizontal and upper left ordinate) and lower right horizontal coordinates of the upper body rectangle of each human body in a video frame remain unchanged compared to the human body rectangle coordinates of that human body in that video frame. For example, 1) Set the human body keypoint score threshold to score_threshold=0.3;
[0044] 2) Let the coordinates of the top left and bottom right corners of the human detection rectangle be (x1, y1) and (x2, y2) respectively.
[0045] 3) Record the coordinates and scores of multiple key human body points: left shoulder, right shoulder, left hip, and right hip.
[0046] a.left_shoulder_x,left_shoulder_y,left_shoulder_score;
[0047] b.right_shoulder_x,right_shoulder_y,right_shoulder_score;
[0048] c.left_hip_x, left_hip_y, left_hip_score;
[0049] d.right_hip_x, right_hip_y, right_hip_score;
[0050] 4) Calculate avg_shoulder_y (average shoulder ordinate), and set its initial value to None;
[0051] a. If left_shoulder_score is greater than score_threshold and right_shoulder_score is greater than score_threshold, then avg_shoulder_y is equal to (left_shoulder_y + right_shoulder_y) / 2;
[0052] b. If left_shoulder_score is greater than score_threshold and right_shoulder_score is less than score_threshold, then denote avg_shoulder_y as equal to left_shoulder_y;
[0053] c. If left_shoulder_score is less than score_threshold and right_shoulder_score is greater than score_threshold, then denote avg_shoulder_y as equal to right_shoulder_y;
[0054] 5) Calculate upper_body_bottom (coordinates of the lower boundary of the upper body), and set its initial value to None;
[0055] a. If left_hip_score is greater than score_threshold and right_hip_score is greater than score_threshold, then set upper_body_bottom as max(left_hip_y, right_hip_y);
[0056] b. If left_hip_score is greater than score_threshold and right_hip_score is less than score_threshold, then denote upper_body_bottom as equal to left_hip_y;
[0057] c. If left_hip_score is less than score_threshold and right_hip_score is greater than score_threshold, then denote upper_body_bottom as equal to right_hip_y;
[0058] d. If left_hip_score is less than score_threshold and right_hip_score is less than score_threshold, then
[0059] a) If avg_shoulder_y is not equal to None, then upper_body_bottom equals avg_shoulder_y + (y² - avg_shoulder_y) * 0.7;
[0060] 6) Calculate upper_y2 (the y-coordinate of the bottom right corner of the upper body rectangle);
[0061] a. If upper_body_bottom equals None, then upper_y2 equals y1 + (y2 - y1) * 0.6;
[0062] b. If upper_body_bottom is not equal to None
[0063] a) Calculate min_height, which is equal to (y2 - y1) * 0.4;
[0064] b) upper_y2 equals max(min(upper_body_bottom, y2), y1 + min_height);
[0065] 7) Return the coordinates of the upper body rectangle, with the top left corner at (x1, y1) and the bottom right corner at (x2, upper_y2).
[0066] In some embodiments, the step of performing an affine transformation on each video frame based on the upper body bounding box coordinates to obtain the upper body image frame corresponding to each video frame includes: constructing a corresponding affine matrix based on the upper body bounding box coordinates and the target image size; performing an affine transformation on each video frame based on the affine matrix to obtain the upper body image frame corresponding to each video frame. In some embodiments, for the upper body bounding box coordinates corresponding to each human body in each video frame contained in the video stream, a 2*3 affine matrix is constructed based on the upper body bounding box coordinates and the target image size (e.g., 224*224); an affine transformation (e.g., bilinear interpolation affine transformation) is performed on the video frame based on the affine matrix to obtain the upper body image frame corresponding to the human body in that video frame, wherein the size of the upper body image frame is the target image size.
[0067] In some embodiments, the method further includes: converting the coordinates of the upper body bounding box into a format that includes center point coordinates and a scale format; wherein, constructing the corresponding affine matrix based on the upper body bounding box coordinates and the target image size includes: constructing the corresponding affine matrix based on the converted upper body bounding box coordinates and the target image size. In some embodiments, the upper body bounding box coordinates are in a format that includes the coordinates of the upper left corner and the lower right corner. Therefore, the upper body bounding box coordinates need to be converted into a format that includes the center point coordinates and the scale (length and width), and then a 2*3 affine matrix is constructed based on the converted upper body bounding box coordinates and the target image size.
[0068] In some embodiments, the method further includes: expanding the upper body bounding box coordinates according to preset initialization parameters to obtain expanded upper body bounding box coordinates; wherein, constructing a corresponding affine matrix based on the upper body bounding box coordinates and the target image size includes: constructing a corresponding affine matrix based on the expanded upper body bounding box coordinates and the target image size. In some embodiments, the upper body bounding box coordinates need to be expanded first according to a preset initialization parameter scale_rate (e.g., 1.25) to ensure that the joints are not truncated, and then a 2*3 affine matrix is constructed based on the expanded upper body bounding box coordinates and the target image size. In some embodiments, the upper body bounding box coordinates can be expanded first, then the expanded upper body bounding box coordinates can be format-converted, and then a 2*3 affine matrix can be constructed based on the format-converted upper body bounding box coordinates and the target image size.
[0069] In some embodiments, obtaining a corresponding dynamic gesture recognition result by performing dynamic gesture recognition on the upper body image frame includes: inputting the upper body image frame into a lightweight convolutional neural network model to obtain a dynamic gesture recognition result corresponding to the upper body image frame, wherein the modules with a preset number of layers in the lightweight convolutional neural network model are replaced with inverse residual modules with channel shifting. In some embodiments, the upper body image frame is input into an improved lightweight convolutional neural network model to obtain the dynamic gesture recognition result corresponding to the upper body image frame output by the model. In some embodiments, a lightweight convolutional neural network model (e.g., the MobileNetV2 model) is improved by replacing the InvertedResidual Module with a preset number of layers (e.g., 3, 5, 6, 8, 9, 10, 12, 13, 15, and 16 layers) with an InvertedResidualWithChannelShift Module. This allows the temporal information to be "moved" along the channel dimension of the 2D convolutional neural network, giving the convolutional kernel a 3D receptive field, while increasing the number of parameters and computational cost by almost nothing. In low-latency streaming inference, the convolutional kernel can see both the previous and current frames simultaneously, doubling the temporal receptive field while maintaining a 2D computational cost. Compared to the InvertedResidual Module, the InvertedResidualWithChannelShift Module does not introduce additional parameters or change the network structure; it only achieves temporal information fusion through data rearrangement along the channel dimension, thus improving model performance.
[0070] In some embodiments, obtaining the dynamic gesture recognition result corresponding to the upper body image frame by inputting the upper body image frame into a lightweight convolutional neural network model includes: merging the upper body image frame with the current gesture buffer data to obtain corresponding input information; inputting the input information into the lightweight convolutional neural network model, causing the lightweight convolutional neural network model to perform model inference on the input information to obtain the logical features corresponding to the upper body image frame output by the lightweight convolutional neural network model and the updated gesture buffer data; obtaining the average logical features corresponding to the upper body image frame based on the logical features corresponding to the upper body image frame and the logical features corresponding to other upper body image frames preceding the upper body image frame during the accumulation process; normalizing the average logical features to obtain the dynamic gesture recognition result corresponding to the upper body image frame. In some embodiments, the gesture buffer data includes ten elements, for example, the width and height of the upper body image frame are denoted as Width and Height, and the initial values of the gesture buffer data are as follows:
[0071] 1) The first element has a shape (i.e., shape, the core property describing the size of the tensor's dimensions, which is essentially a tuple / list that quantifies the number of elements in each dimension of the tensor) of (1, 3, Height / 4, Width / 4), all of which have a value of 0;
[0072] 2) The second element has a shape of (1, 4, Height / 8, Width / 8) and all values are 0;
[0073] 3) The third element has a shape of (1, 4, Height / 8, Width / 8) and all values are 0;
[0074] 4) The fourth element has a shape of (1, 8, Height / 16, Width / 16) and all values are 0;
[0075] 5) The fifth element has a shape of (1, 8, Height / 16, Width / 16) and all values are 0;
[0076] 6) The sixth element has a shape of (1, 8, Height / 16, Width / 16) and all values are 0;
[0077] 7) The seventh element has a shape of (1, 12, Height / 16, Width / 16) and all values are 0;
[0078] 8) The eighth element has a shape of (1, 12, Height / 16, Width / 16) and all values are 0;
[0079] 9) The ninth element has a shape of (1, 20, Height / 32, Width / 32) and all values are 0;
[0080] 10) The tenth element has a shape of (1, 20, Height / 32, Width / 32) and all values are 0;
[0081] Each element has a shape format of (1, channel, height, width). For each upper body image frame, it is merged with the current gesture buffer data (the gesture buffer data after the previous update, or the gesture buffer data after initialization) to obtain the corresponding input information. For example, the length of the upper body image frame is 1, and its shape is (1, 3, height, width). The gesture buffer data includes ten elements, and the shape of each element is as described above. The data length of the merged input information is 11. The first element of the input information is the upper body image frame, and the second to eleventh elements of the input information are the gesture buffer data. In some embodiments, an improved lightweight convolutional neural network model is used to perform model inference based on the input information to obtain the logical features corresponding to the upper body image frame and the updated gesture buffer data (i.e., updating the current gesture buffer data to obtain the updated gesture buffer data). The shape of the logical features is (1,2), where 2 is the number of categories (wake-up and unwake-up). Then, based on the logical features corresponding to the upper body image frame and the logical features corresponding to other upper body image frames preceding the upper body image frame during the accumulation process, the average of these multiple logical features is calculated to obtain the average logical features corresponding to the upper body image frame. Then, a softmax normalization operation is performed on the average logical features. The dynamic gesture recognition result corresponding to the upper body image frame is determined based on the normalized average logical features. For example, if the normalized average logical features are 1, the dynamic gesture recognition result corresponding to the upper body image frame is "wake-up"; if the normalized average logical features are 0, the dynamic gesture recognition result corresponding to the upper body image frame is "unwake-up". Here, softmax is the core activation function, whose core function is to output any real number (such as the model's logits). The score is transformed into a probability distribution in the 0-1 interval, and the sum of all output values is 1. Its essence is a normalization tool, which not only preserves the relative magnitude of the original output, but also directly maps it to an interpretable probability.
[0082] In some embodiments, the method further includes: standardizing the upper body image frame based on the mean and standard deviation of the images in a preset image dataset with respect to the three channels to obtain a standardized upper body image frame; wherein, merging the upper body image frame with the current gesture buffer data to obtain corresponding input information includes: merging the standardized upper body image frame with the current gesture buffer data to obtain corresponding input information. In some embodiments, the upper body image frame is first standardized, and then the standardized upper body image frame is merged with the current gesture buffer data. The standardization process refers to standardizing the upper body image frame according to the mean and standard layer calculated in the preset image dataset (e.g., the ImageNet dataset) according to the three RGB channels, that is, subtracting the mean and dividing by the standard deviation, and the calculation formula is img_norm = (img – mean) / std. Wherein, img is the upper body image frame before standardization (the pixel values of the RGB three channels are all divided by 255), mean represents the mean of the images in ImageNet (e.g., 0.406, 0.456, 0.485), std represents the standard deviation of the images in ImageNet (e.g., 0.225, 0.224, 0.229), and img_norm is the upper body image frame after standardization. The standardized data can accelerate model convergence during training and effectively improve the model's generalization ability.
[0083] In some embodiments, performing a wake-up operation on the robot further includes clearing the upper body image frames of the video stream related to the at least one human body. In some embodiments, after performing a wake-up operation on the robot, the upper body image frames of each currently playing video frame related to each human body included in the video stream are cleared, ending the current wake-up process.
[0084] In some embodiments, the method further includes: if the number of upper-body image frames in the plurality of upper-body image frames that detect a preset gesture does not meet the preset condition, deleting the first upper-body image frame in the plurality of upper-body image frames, and waiting for the plurality of upper-body image frames to accumulate again to reach the first preset number of frames. In some embodiments, if the number of upper-body image frames in the plurality of upper-body image frames corresponding to a certain human body (e.g., 16 upper-body image frames, frames 1-16) that detect a preset gesture does not meet the preset condition, for example, if the number of upper-body image frames that detect a preset gesture is less than a second preset number of frames (e.g., 8 frames) or the ratio of the number of upper-body image frames that detect a preset gesture to the total number of frames in the currently accumulated plurality of upper-body image frames is less than a preset ratio threshold, then deleting the first upper-body image frame in the plurality of upper-body image frames. The robot first accumulates upper body image frames (e.g., frame 1), and then waits for the corresponding upper body image frames to accumulate again to reach the first preset number of frames (e.g., frames 2-17). If the dynamic gesture recognition result corresponding to the current upper body image frame (e.g., frame 17) is wake-up, static gesture detection is performed on the currently accumulated upper body image frames (e.g., frames 2-17). If the number of upper body image frames in the multiple upper body image frames that have detected preset gestures is greater than or equal to the second preset number of frames, a wake-up operation is performed on the robot, and so on.
[0085] Figure 2 A flowchart of a method for dynamic gesture recognition according to an embodiment of this application is shown.
[0086] like Figure 2 As shown, the input to the dynamic gesture recognition module is a sequence of upper body image frames (including 16 frames) consisting of upper body image frames corresponding to a certain human body in each video frame contained in the video stream. The output of the dynamic gesture recognition module includes two results: wake-up and non-wake-up.
[0087] Figure 3 A flowchart of a method for static gesture recognition according to an embodiment of this application is shown.
[0088] like Figure 3 As shown, the input to the static gesture recognition module is each upper body image frame from the multiple upper body image frames currently accumulated for a certain human body contained in the video stream, and the output of the static gesture recognition module is the palm gesture rectangle detection box (e.g., Figure 3 The coordinates of the rectangle in the image and the category probability corresponding to the palm gesture (e.g., ...). Figure 3 (0.97 in the middle).
[0089] Figure 4The diagram illustrates a computer device structure for waking up a robot according to an embodiment of this application. The computer device includes a first module 11, a second module 12, and a third module 13. The first module 11 is used to obtain the upper body bounding box coordinates of at least one human body in each video frame of the video stream obtained by the robot. The second module 12 is used to perform an affine transformation on each video frame based on the upper body bounding box coordinates to obtain an upper body image frame of each video frame with respect to the at least one human body. The third module 13 is used to perform dynamic gesture recognition on the upper body image frames to obtain corresponding dynamic gesture recognition results. If the accumulated number of upper body image frames reaches a first preset frame number, and the dynamic gesture recognition result corresponding to the current upper body image frame is "wake-up," static gesture detection is performed on the currently accumulated multiple upper body image frames. If the number of upper body image frames in the multiple upper body image frames that have detected a preset gesture meets a preset condition, a wake-up operation is performed on the robot.
[0090] Module 11 is used to obtain the coordinates of the upper body human body rectangle corresponding to at least one human body in each video frame of the video stream obtained by the robot.
[0091] In some embodiments, the computer device may be the robot itself, or it may be a remote workstation responsible for model reasoning.
[0092] In some embodiments, the robot performs target (here, the target only includes human bodies) tracking on the video stream obtained by its camera, obtaining the identification information (e.g., ID) of at least one human body contained in each video frame of the video stream, and the coordinates of the bounding box corresponding to the upper body of each human body in each video frame, i.e., the coordinates of the upper body human body bounding box. In some embodiments, the upper body human body bounding box coordinates include the coordinates of the four vertices of the upper body human body bounding box, or, only the coordinates of the two diagonal vertices of the upper body human body bounding box, for example, only the coordinates of the upper left corner and the lower right corner, or only the coordinates of the lower left corner and the upper right corner.
[0093] Module 12 is used to perform an affine transformation on each video frame based on the coordinates of the upper body human body rectangle to obtain an upper body image frame of each video frame with respect to at least one human body.
[0094] In some embodiments, an affine transformation is a linear geometric transformation that preserves the parallel lines and the proportions of collinear points. It combines scaling, rotation, cropping, and translation effects by applying linear transformations and translation operations to the vertices of an image or graphic. In some embodiments, for each video frame in a video stream, an affine transformation (e.g., bilinear interpolation affine transformation) is performed on the video frame based on the coordinates of the upper body bounding box corresponding to each human figure in the video frame. This extracts the image of the ROI (Region of Interest) for each human figure in the video frame, i.e., the upper body image frame. By obtaining the upper body image frame of each human figure in each video frame through the affine transformation, the image area containing the upper body bounding box of each human figure in the video frame is cropped and scaled to a fixed size while maintaining the proportions of the human figures (avoiding image distortion and keeping the target in the original image from deforming).
[0095] Module 13 is used to perform dynamic gesture recognition on the upper body image frames to obtain the corresponding dynamic gesture recognition results. If the cumulative number of upper body image frames reaches a first preset number of frames and the dynamic gesture recognition result corresponding to the current upper body image frame is wake-up, static gesture detection is performed on the currently accumulated multiple upper body image frames. If the number of upper body image frames in the multiple upper body image frames that have been detected with preset gestures meets the preset conditions, a wake-up operation is performed on the robot.
[0096] In some embodiments, for each human body contained in the video stream, dynamic gesture recognition is performed on the upper body image frame corresponding to that human body in each video frame to obtain the dynamic gesture recognition result corresponding to that upper body image frame. Here, dynamic gesture recognition is a technology that captures, analyzes, and semantically interprets the motion trajectory, action sequence, and posture change process of a user's hand in a continuous temporal space. The core is to recognize gesture actions with a time dimension, rather than a single fixed posture. It recognizes the continuous motion trajectory and temporal change process of the gesture, which requires analyzing the position and posture evolution law of the hand in multiple consecutive images. Essentially, it is the analysis of a temporal video sequence, that is, the analysis of the upper body image frame sequence composed of the upper body image frames corresponding to each video frame in the video stream.
[0097] In some embodiments, for each human body included in the video stream, the dynamic gesture recognition result corresponding to each upper body image frame in the upper body image frame sequence corresponding to that human body in each video frame can be obtained by inputting the upper body image frame sequence (e.g., MobileNetV2 model, a lightweight convolutional neural network proposed by Google, designed for mobile and embedded devices, and an upgrade of MobileNetV1) into a lightweight convolutional neural network model. The lightweight convolutional neural network model is a type of convolutional neural network (CNN) designed for computationally limited scenarios (such as mobile devices, embedded devices, and edge computing terminals). Its core objective is to achieve fast inference and low resource consumption by simplifying the network structure, reducing the number of parameters and computational load while maintaining a certain level of accuracy. In some embodiments, the dynamic gesture recognition result includes both awake and non-awakened results.
[0098] In some embodiments, for each human body included in the video stream, if the cumulative number of upper body image frames corresponding to that human body reaches a first preset number of frames (e.g., 16 frames), and the dynamic gesture recognition result of the current upper body image frame (e.g., the 16th frame) corresponding to that human body is "wake-up," then static gesture detection is performed on each of the current cumulative upper body image frames (i.e., the 16 upper body image frames). For each upper body image frame, if a preset gesture is recognized from that upper body image frame, the static gesture recognition result corresponding to that upper body image frame is output. The static gesture recognition result includes the coordinates of the rectangular detection box corresponding to the preset gesture in the upper body image frame and the category probability corresponding to the preset gesture. Based on the static gesture recognition result, it can be determined whether the preset gesture has been detected in the upper body image frame. Static gesture detection is a recognition technology for gestures in a static state. It focuses on the spatial morphological features of the gesture, rather than the temporal change features of the action. The object of static gesture detection is the gesture posture at a certain moment, such as an open palm, a clenched fist, or making an "OK" sign. Fixed gestures, such as raising the index finger, can be detected and recognized using only a single image frame, without requiring a continuous sequence of video frames. The core basis for recognition is the spatial geometric features of the gesture, including its shape and outline, the number of fingers, joint positions, and palm orientation. In some embodiments, the yolov5n model (an extremely lightweight version of the YOLOv5 object detection algorithm family developed by the Ultralytics team, designed specifically for edge devices, mobile devices, and low-power hardware with limited computing power; it is the model with the fewest parameters and the fastest inference speed in the YOLOv5 series) can be used to perform static gesture detection on upper body image frames. That is, the upper body image frame is input into the yolov5n model, and the static gesture recognition result corresponding to the upper body image frame is obtained from the network's output. In some embodiments, the rule module judges the static gesture recognition results corresponding to the currently accumulated multiple upper body image frames. If the number of upper body image frames (e.g., 16 upper body image frames) that detect a preset gesture meets a preset condition, it is determined that the robot needs to be woken up and the corresponding wake-up operation is performed on the robot. Otherwise, it is determined that the robot does not need to be woken up. The preset condition includes, but is not limited to, the number of upper body image frames that detect a preset gesture being greater than or equal to a second preset number of frames (e.g., 8 frames), and the ratio of the number of upper body image frames that detect a preset gesture to the total number of frames of the currently accumulated multiple upper body image frames being greater than or equal to a preset ratio threshold. This example embodiment does not impose any special limitations on these conditions. In some embodiments, the rule module can expand and update the rules according to the usage scenario. In some embodiments, at a frame rate of 30fps, 16 frames are approximately equal to 0.5-0.7 seconds of video, which is sufficient to cover a complete cycle of most waving actions, ensuring that the model can capture the start and end of the action, and achieving a trade-off between the temporal receptive field and computational efficiency.In some embodiments, performing a wake-up operation on a robot refers to the process of activating the robot from a dormant or low-power state, enabling it to enter an interactive working state.
[0099] In some embodiments, obtaining the upper body bounding box coordinates of at least one human body in each video frame of the video stream obtained by the robot includes: performing pedestrian tracking on the video stream obtained by the robot to obtain the human bounding box coordinates and their corresponding first score for at least one human body in each video frame, and the coordinates of multiple key points corresponding to the at least one human body and their corresponding second score; determining the upper body bounding box coordinates of at least one human body in each video frame based on the human bounding box coordinates, the first score, the key point coordinates, and the second score. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0100] In some embodiments, determining the upper body rectangle coordinates of at least one human body in each video frame based on the human body rectangle coordinates, the first score, the keypoint coordinates, and the second score includes: determining the lower right corner ordinate of the upper body rectangle of at least one human body in each video frame based on the lower right corner ordinate of the human body rectangle, the first score, the keypoint coordinates, and the second score, wherein the upper left corner coordinates and lower right corner x-coordinates of the upper body rectangle remain unchanged compared to the human body rectangle coordinates. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0101] In some embodiments, the step of performing an affine transformation on each video frame based on the coordinates of the upper body bounding box to obtain the upper body image frame corresponding to each video frame includes: constructing a corresponding affine matrix based on the coordinates of the upper body bounding box and the target image size; and performing an affine transformation on each video frame based on the affine matrix to obtain the upper body image frame corresponding to each video frame. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0102] In some embodiments, the computer device is further configured to: perform format conversion on the upper body human body rectangle coordinates, converting the upper body human body rectangle coordinates into center point coordinates and scale format; wherein, the step of constructing a corresponding affine matrix based on the upper body human body rectangle coordinates and the target image size includes: constructing a corresponding affine matrix based on the format-converted upper body human body rectangle coordinates and the target image size. Here, the related operations are similar to... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0103] In some embodiments, the computer device is further configured to: expand the coordinates of the upper body human body rectangle according to preset initialization parameters to obtain the expanded upper body human body rectangle coordinates; wherein, the step of constructing a corresponding affine matrix based on the upper body human body rectangle coordinates and the target image size includes: constructing a corresponding affine matrix based on the expanded upper body human body rectangle coordinates and the target image size. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0104] In some embodiments, the step of performing dynamic gesture recognition on the upper body image frame to obtain the corresponding dynamic gesture recognition result includes: inputting the upper body image frame into a lightweight convolutional neural network model to obtain the dynamic gesture recognition result corresponding to the upper body image frame, wherein the module with a preset number of layers in the lightweight convolutional neural network model is replaced with an inverted residual module with channel shifting.
[0105] In some embodiments, obtaining the dynamic gesture recognition result corresponding to the upper body image frame by inputting the upper body image frame into a lightweight convolutional neural network model includes: merging the upper body image frame with the current gesture buffer data to obtain corresponding input information; inputting the input information into the lightweight convolutional neural network model, causing the lightweight convolutional neural network model to perform model inference on the input information to obtain the logical features corresponding to the upper body image frame output by the lightweight convolutional neural network model and the updated gesture buffer data; obtaining the average logical features corresponding to the upper body image frame based on the logical features corresponding to the upper body image frame and the logical features corresponding to other upper body image frames preceding the upper body image frame during the accumulation process; performing a normalization operation on the average logical features to obtain the dynamic gesture recognition result corresponding to the upper body image frame. Here, the related operations are similar to... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0106] In some embodiments, the computer device is further configured to: standardize the upper body image frame based on the mean and standard deviation of the images in a preset image dataset with respect to the three channels, to obtain a standardized upper body image frame; wherein, merging the upper body image frame with the current gesture buffer data to obtain corresponding input information includes: merging the standardized upper body image frame with the current gesture buffer data to obtain corresponding input information. Here, related operations are... Figure 1The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0107] In some embodiments, performing the wake-up operation on the robot further includes: clearing the video stream of upper body image frames related to the at least one human body. Here, the related operation is... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0108] In some embodiments, the computer device is further configured to: if the number of upper-body image frames in the plurality of upper-body image frames that detect a preset gesture does not meet the preset condition, delete the first upper-body image frame in the plurality of upper-body image frames, and wait for the plurality of upper-body image frames to accumulate again to reach the first preset number of frames. Here, the related operations are... Figure 1 The embodiments shown are the same or similar, so they will not be described again, but are included here by reference.
[0109] Figure 5 Exemplary systems that can be used to implement the various embodiments described in this application are shown; such as Figure 5 As shown in some embodiments, system 300 can function as any of the devices described in each of the embodiments. In some embodiments, system 300 may include one or more computer-readable media having instructions (e.g., system memory or NVM / storage device 320) and one or more processors (e.g., one or more processors 305) coupled to the one or more computer-readable media and configured to execute the instructions to implement the module and thus perform the actions described in this application.
[0110] In one embodiment, the system control module 310 may include any suitable interface controller to provide any suitable interface to at least one of the processors 305 and / or any suitable device or component communicating with the system control module 310.
[0111] The system control module 310 may include a memory controller module 330 to provide an interface to the system memory 315. The memory controller module 330 may be a hardware module, a software module, and / or a firmware module.
[0112] System memory 315 can be used, for example, to load and store data and / or instructions for system 300. In one embodiment, system memory 315 may include any suitable volatile memory, such as suitable DRAM. In some embodiments, system memory 315 may include double data rate type quad synchronous dynamic random access memory (DDR4 SDRAM).
[0113] In one embodiment, the system control module 310 may include one or more input / output (I / O) controllers to provide interfaces to the NVM / storage device 320 and (one or more) communication interfaces 325.
[0114] For example, NVM / storage device 320 may be used to store data and / or instructions. NVM / storage device 320 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drive (HDD), one or more optical disc (CD) drives, and / or one or more digital universal optical disc (DVD) drives).
[0115] NVM / storage device 320 may include storage resources that are physically part of a device on which system 300 is mounted, or that can be accessed by the device without necessarily being part of it. For example, NVM / storage device 320 may be accessed via a network through one or more communication interfaces 325.
[0116] One or more communication interfaces 325 may provide the system 300 with an interface to communicate over one or more networks and / or with any other suitable device. The system 300 may wirelessly communicate with one or more components of a wireless network in accordance with any of one or more wireless network standards and / or protocols.
[0117] In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 (e.g., memory controller module 330). In one embodiment, at least one of the processors 305 may be logically packaged with one or more controllers of the system control module 310 to form a system-in-package (SiP). In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die. In one embodiment, at least one of the processors 305 may be integrated with the logic of one or more controllers of the system control module 310 on the same die to form a system-on-a-chip (SoC).
[0118] In various embodiments, system 300 may be, but is not limited to, a server, workstation, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, system 300 may have more or fewer components and / or different architectures. For example, in some embodiments, system 300 includes one or more cameras, a keyboard, a liquid crystal display (LCD) screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.
[0119] In addition to the methods and devices described in the above embodiments, this application also provides a computer-readable storage medium storing computer code that, when executed, performs the method described in any of the preceding embodiments.
[0120] This application also provides a computer program product that, when executed by a computer device, performs the method described in any of the preceding claims.
[0121] This application also provides a computer device, the computer device comprising:
[0122] One or more processors;
[0123] Memory, used to store one or more computer programs;
[0124] When the one or more computer programs are executed by the one or more processors, the one or more processors cause the one or more processors to perform the method as described in any of the preceding methods.
[0125] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0126] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. Those skilled in the art will understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executing the instructions, or the computer compiling the instructions and then executing the corresponding compiled program, or the computer reading and executing the instructions, or the computer reading and installing the instructions and then executing the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to a computer.
[0127] Communication media include media through which communication signals containing, for example, computer-readable instructions, data structures, program modules, or other data are transmitted from one system to another. Communication media can include guided transmission media (such as cables and wires (e.g., optical fibers, coaxial cables, etc.)) and wireless (unguided transmission) media capable of propagating energy waves, such as sound, electromagnetic, RF, microwave, and infrared. Computer-readable instructions, data structures, program modules, or other data can be embodied as modulated data signals in, for example, wireless media (such as carrier waves or similar mechanisms embodied as part of spread spectrum technology). The term "modulated data signal" refers to a signal whose one or more characteristics are altered or set in a manner that encodes information in the signal. Modulation can be analog, digital, or a hybrid modulation technique.
[0128] By way of example and not limitation, computer-readable storage media may include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. For example, computer-readable storage media include, but are not limited to, volatile memories such as random access memory (RAM, DRAM, SRAM); and non-volatile memories such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic and ferromagnetic / ferroelectric memories (MRAM, FeRAM); and magnetic and optical storage devices (hard disks, magnetic tapes, CDs, DVDs); or other media now known or hereafter developed capable of storing computer-readable information / data for use by a computer system.
[0129] Herein, one embodiment of this application includes an apparatus comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein when the computer program instructions are executed by the processor, the apparatus is triggered to run a method and / or technical solution based on the foregoing embodiments of this application.
[0130] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.
Claims
1. A method for waking up a robot, wherein, The method includes: Based on the video stream obtained by the robot, obtain the coordinates of the upper body human body rectangle corresponding to at least one human body in each video frame of the video stream; Affine transformation is performed on each video frame based on the coordinates of the upper body human body rectangle to obtain an upper body image frame of each video frame with respect to the at least one human body. The upper body image frame is merged with the current gesture buffer data to obtain the corresponding input information; The input information is input into a lightweight convolutional neural network model, which then performs model inference on the input information to obtain the logical features corresponding to the upper body image frame and the updated gesture buffer data output by the lightweight convolutional neural network model. The modules with a preset number of layers in the lightweight convolutional neural network model are replaced with inverse residual modules with channel shift. Based on the logical features corresponding to the upper body image frame and the logical features corresponding to other upper body image frames preceding the upper body image frame during the accumulation process, the average logical features corresponding to the upper body image frame are obtained. The average logical features are then normalized to obtain the dynamic gesture recognition result corresponding to the upper body image frame. If the cumulative number of upper body image frames reaches the first preset number of frames, and the dynamic gesture recognition result corresponding to the current upper body image frame is wake-up, static gesture detection is performed on the currently accumulated multiple upper body image frames. If the number of upper body image frames in the multiple upper body image frames that have been detected with preset gestures meets the preset conditions, the robot is woken up.
2. The method according to claim 1, wherein, The step of obtaining the upper body bounding box coordinates of at least one human body in each video frame of the video stream obtained by the robot includes: The robot performs pedestrian tracking on the video stream it obtains, and obtains the human body bounding box coordinates and its corresponding first score for at least one human body in each video frame of the video stream, as well as the coordinates of multiple key points corresponding to the at least one human body and their corresponding second scores. Based on the human body bounding box coordinates, the first score, the key point coordinates, and the second score, determine the upper body bounding box coordinates corresponding to at least one human body in each video frame.
3. The method according to claim 2, wherein, The step of determining the upper body bounding box coordinates of at least one human body in each video frame based on the human body bounding box coordinates, the first score, the key point coordinates, and the second score includes: Based on the lower right corner ordinate of the human body rectangle, the first score, the key point coordinates, and the second score, the lower right corner ordinate of the upper body human body rectangle corresponding to at least one human body in each video frame is determined, wherein the upper left corner coordinate and lower right corner x-coordinate of the upper body human body rectangle remain unchanged compared to the human body rectangle coordinates.
4. The method according to claim 1, wherein, The step of performing an affine transformation on each video frame based on the coordinates of the upper body human body rectangle to obtain the upper body image frame corresponding to each video frame includes: Based on the coordinates of the upper body human body rectangle and the size of the target image, a corresponding affine matrix is constructed. Based on the affine matrix, an affine transformation is performed on each video frame to obtain the upper body image frame corresponding to each video frame.
5. The method according to claim 4, wherein, The method further includes: The coordinates of the upper body human body rectangle are converted into center point coordinates and scale format. The step of constructing the corresponding affine matrix based on the coordinates of the upper body bounding box and the target image size includes: Construct the corresponding affine matrix based on the coordinates of the upper body rectangle after format conversion and the size of the target image.
6. The method according to claim 4, wherein, The method further includes: The coordinates of the upper body human body rectangle are expanded outward according to the preset initialization parameters to obtain the coordinates of the expanded upper body human body rectangle. The step of constructing the corresponding affine matrix based on the coordinates of the upper body bounding box and the target image size includes: Construct a corresponding affine matrix based on the coordinates of the expanded upper body rectangle and the size of the target image.
7. The method according to claim 1, wherein, The method further includes: Based on the mean and standard deviation of the images in the preset image dataset with respect to the three channels, the upper body image frames are standardized to obtain standardized upper body image frames. The step of merging the upper body image frame with the current gesture buffer data to obtain the corresponding input information includes: The standardized upper body image frame is merged with the current gesture buffer data to obtain the corresponding input information.
8. The method according to claim 1, wherein, The step of waking up the robot further includes: Clean up the video stream of the upper body image frames of the at least one human body.
9. The method according to claim 1, wherein, The method further includes: If the number of upper body image frames in the plurality of upper body image frames that detect the preset gesture does not meet the preset condition, delete the first upper body image frame in the plurality of upper body image frames, and wait for the plurality of upper body image frames to accumulate again to reach the first preset number of frames.
10. A computer device for waking up a robot, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 9.
11. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 9.
12. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 9.