Robot control method and system based on human body posture detection
By combining image sensors and three-dimensional data sensors to collect data, and using feature extraction and classification algorithm models for human posture recognition, the problem of insufficient human posture recognition accuracy of existing teaching robots in complex environments is solved, and high-precision and low-cost personalized teaching interaction is achieved.
Patent Information
- Application Number
- CN202510507426.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-04
AI Technical Summary
The existing teaching robots have insufficient human posture recognition accuracy in high-precision operation and complex environments, which are greatly affected by environmental conditions, and the high-precision equipment costs, which limits their wide application in educational scenarios.
By combining image sensors and three-dimensional data sensors to collect two-dimensional and three-dimensional data, preset feature extraction rules and classification algorithm models are used to perform feature fusion and pose recognition, and action intentions are obtained to control the robot to perform interactive actions and adapt to different environments and scenarios.
It realizes accurate detection of human posture in complex environments, reduces environmental interference, provides personalized teaching content, reduces equipment costs, and improves user experience and applicability.
Smart Images

Figure CN120244969A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of teaching robots, and in particular, to a robot control method and system based on human body posture detection. Background Art
[0002] With the continuous development of artificial intelligence, machine learning, sensor technology, etc., it provides technical support for the performance improvement and function expansion of teaching robots.
[0003] Currently, devices such as cameras and depth sensors are usually used to comprehensively apply various recognition technologies such as voice, gesture, and vision to understand human instructions. In the actual teaching environment, there may be various interference factors, such as noise, light changes, multiple people giving instructions simultaneously, etc., which will affect the accurate recognition of action instructions by the robot. For some fine action instructions that require high-precision operations, the current technology may still be difficult to accurately recognize and execute. And relatively high-precision devices (such as wearable exoskeleton devices) are often expensive, which may pose financial pressure on some schools and educational institutions and limit their large-scale popularization and application. Summary of the Invention
[0004] In order to solve the above problems, embodiments of the present application provide a robot control method and system, an electronic device, a computer-readable storage medium, and a computer program product based on human body posture detection.
[0005] In a first aspect, the present application provides a robot control method based on human body posture detection, including:
[0006] Obtaining human body posture data of a target object in a target scene, where the human body posture data includes two-dimensional data and three-dimensional data;
[0007] Performing feature extraction on the human body posture data based on a preset feature extraction rule to obtain key feature data corresponding to the target object, where the key feature data includes at least one type of human body posture feature corresponding to the target object;
[0008] Performing feature fusion processing on the key feature data to obtain a comprehensive feature vector;
[0009] Performing posture recognition on the comprehensive feature vector based on a preset classification algorithm model to obtain a posture recognition result;
[0010] Obtaining the correspondence between a preset posture recognition result and an action intention in the target scene, and based on the correspondence and the posture recognition result, obtaining a target action intention matching the posture recognition result;
[0011] Based on the posture recognition result and the target action intention, a control instruction is obtained to control the robot to execute corresponding interaction actions based on the control instruction.
[0012] The beneficial effects are as follows:
[0013] In the technical solution provided by the embodiment of the present application, an image sensor and a three-dimensional data sensor are used to periodically collect the human body posture data of the target object in the target scene, including two-dimensional data and three-dimensional data. Based on a preset feature extraction rule, feature extraction is performed on the two-dimensional data and the three-dimensional data to obtain corresponding key feature data, and the key feature data includes at least one type of human body posture feature corresponding to the target object. Then, feature fusion processing is performed on the key feature data to obtain a comprehensive feature vector; based on a preset classification algorithm model, posture recognition is performed on the comprehensive feature vector to obtain a posture recognition result; the corresponding relationship between the preset posture recognition result and the action intention in the target scene is obtained, and based on the corresponding relationship and the posture recognition result, a target action intention matching the posture recognition result is obtained. Finally, a control instruction is obtained based on the posture recognition result and the target action intention, and the robot is controlled to execute corresponding interaction actions based on the control instruction. In this way, the present application can real-time sense the actions and postures of people in the education scene, through the fusion of two-dimensional data and three-dimensional data, while realizing accurate human body posture detection, reducing the influence of the environment on posture recognition, to adapt to different target scenes and environmental conditions, and according to the corresponding relationship between the preset posture recognition result and the action intention in the target scene, provide personalized teaching content by controlling the robot.
[0014] In a second aspect, the present invention provides a robot control system based on human body posture detection, including a collection unit, a feature extraction unit, a data fusion unit, an identification unit, a matching unit, and a processing unit;
[0015] The collection unit is used to obtain the human body posture data of the target object in the target scene, and the human body posture data includes two-dimensional data and three-dimensional data;
[0016] The feature extraction unit is used to perform feature extraction on the human body posture data based on a preset feature extraction rule to obtain the key feature data corresponding to the target object, and the key feature data includes at least one type of human body posture feature corresponding to the target object;
[0017] The data fusion unit is used to perform feature fusion processing on the key feature data to obtain a comprehensive feature vector;
[0018] The identification unit is used to perform posture recognition on the comprehensive feature vector based on a preset classification algorithm model to obtain a posture recognition result;
[0019] A matching unit, configured to obtain the correspondence between a preset pose recognition result and an action intention in the target scenario, and based on the correspondence and the pose recognition result, obtain a target action intention that matches the pose recognition result;
[0020] A processing unit, configured to obtain a control instruction based on the pose recognition result and the target action intention, and based on the control instruction, control the robot to execute a corresponding interaction action.
[0021] In a third aspect, the present application further provides an electronic device, including: one or more processors; a storage device, configured to store one or more programs, when the one or more programs are executed by the one or more processors, enabling the electronic device to implement the robot control method based on human pose detection as described above.
[0022] In a fourth aspect, the present application further provides a computer-readable storage medium, on which computer-readable instructions are stored, when the computer-readable instructions are executed by a processor of a computer, enabling the computer to execute the robot control method based on human pose detection as described above.
[0023] In a fifth aspect, the present application further provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the robot control method based on human pose detection provided in the above various alternative embodiments.
[0024] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings
[0025] The drawings here are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:
[0026] Figure 1 is a flowchart of a robot control method based on human pose detection shown in an exemplary embodiment of the present application;
[0027] Figure 2 is a block diagram of a robot control system based on human pose detection shown in an exemplary embodiment of the present application;
[0028] Figure 3 It is a schematic structural diagram of a computer system suitable for an electronic device implementing the embodiments of the present application. Detailed implementation manners
[0029] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0030] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0031] The flowcharts shown in the drawings are only exemplary descriptions and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.
[0032] In the present application, "a plurality of" refers to two or more than two. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0033] In the present application, all actions of obtaining signals, information, or data are carried out on the basis of strictly following the relevant data protection regulations and policies of the country where it is located and with the authorization of the owner of the corresponding device.
[0034] In the related art, the main technology in this field is to collect human body pose data through devices such as cameras and depth sensors, and use computer vision and deep learning algorithms for pose recognition and action classification to achieve human-computer interaction. For example, some educational robots can adjust teaching content or provide feedback by recognizing the limb movements or gestures of the target, enhancing the interactivity and interest of teaching. Or combine cameras and depth sensors to improve the accuracy and robustness of pose detection.
[0035] However, existing systems based on cameras and depth sensors have deficiencies in the accuracy of human pose detection. Although the general pose of the human body can be recognized, in high-precision operation scenarios, such as fine motion recognition or complex interaction tasks, the accuracy limitations are more obvious. Cameras and depth sensors are sensitive to environmental conditions. Under strong light, backlight, or low-light conditions, the quality of the images captured by the camera will be severely affected, resulting in a decrease in the accuracy rate of pose recognition. In addition, depth sensors may not work properly in direct sunlight outdoors. When the human pose is relatively complex, such as when limbs are blocked from each other or the movement amplitude is small, the existing technology may not be able to accurately recognize. For example, when the arm is close to the body, the system may not be able to distinguish between the arm being close to the body and the arm making a small angle with the body. Existing technologies usually require users to operate within a fixed sensor reception range, which limits the user's activity range.
[0036] In addition, although some systems that require wearing exoskeleton devices have high accuracy, they have poor portability and can only be used in fixed locations. In complex environments, the existing technology may make misjudgments, misjudging non-living objects (such as swinging objects) as human actions, causing the robot to make incorrect responses and affecting the teaching effect. Although deep learning-based algorithms can handle complex pose recognition tasks, they have high requirements for computing resources and data volume. In resource-constrained educational scenarios, the real-time performance and reliability of the algorithms may not be guaranteed.
[0037] To solve the above problems, embodiments of the present application propose a robot control method and device, an electronic device, and a computer-readable storage medium based on human pose detection, which mainly involve the robot control technology based on human pose detection included in teaching robots. These embodiments will be described in detail below.
[0038] First, please refer to Figure 1 , Figure 1 which is a flowchart of a robot control method based on human pose detection shown in an exemplary embodiment of the present application. This method can be specifically executed by a server required for the operation of a teaching robot. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. There is no limitation here.
[0039] As Figure 1 shown, in an exemplary embodiment, the robot control method based on human pose detection may include steps S101 to S106, which are introduced in detail as follows:
[0040] Step S101: Obtain the human body pose data of the target object in the target scenario. The human body pose data includes two-dimensional data and three-dimensional data.
[0041] In the embodiments provided by the present application, by periodically and continuously collecting the two-dimensional data and three-dimensional data of the human body pose, real-time pose recognition is achieved. In this way, the two-dimensional data and three-dimensional data are collected at regular time intervals, and after various processes, corresponding control instructions are generated, and then the robot is controlled to execute, ensuring that the robot can respond in a timely manner to the movement changes of the personnel in the target scenario, that is, the progress of the real-time tracking action is realized, and the control instructions are dynamically adjusted. The target scenario includes, but is not limited to, the education scenario.
[0042] Step S102: Extract features from the two-dimensional data and three-dimensional data based on the preset feature extraction rules to obtain the key feature data corresponding to the target object. The key feature data includes at least one type of human body pose feature corresponding to the target object.
[0043] In this embodiment, corresponding preset feature extraction rules are configured for different target scenarios. The preset feature extraction rules set the types of key feature data that need to be recognized in the target scenario, and thus the features of the two-dimensional data and three-dimensional data are extracted to adapt to different target scenarios and environmental conditions.
[0044] Step S103: Perform feature fusion processing on the key feature data to obtain a comprehensive feature vector.
[0045] Step S104: Perform pose recognition on the comprehensive feature vector based on the preset classification algorithm model to obtain a pose recognition result.
[0046] This embodiment combines the two-dimensional data and three-dimensional data, enabling stable recognition of the human body pose under different lighting conditions (such as strong light and weak light), ensuring effective performance in various target scenarios, and reducing the impact of the environment on pose recognition. The comprehensive feature vector obtained after fusing the key feature data is used as a reliable input for the preset classification algorithm model to perform pose recognition on the current human body pose to obtain a pose recognition result.
[0047] Step S105: Obtain the corresponding relationship between the preset pose recognition result and the action intention in the target scenario. Based on the corresponding relationship and the pose recognition result, obtain the target action intention that matches the pose recognition result.
[0048] This application configures the corresponding relationship between the pose recognition result and the action intention for different target scenarios to meet the requirements of the scenario, so that personalized teaching content can be provided according to the postures and action intentions of the users. Therefore, in this embodiment, after the pose recognition result is recognized, based on the corresponding relationship and the pose recognition result, the target action intention that matches the pose recognition result is obtained.
[0049] Step S106, obtain a control instruction based on the pose recognition result and the target action intention, and control the robot to execute the corresponding interaction action based on the control instruction.
[0050] As can be seen from the above, in the method provided in this embodiment, on the one hand, by collecting the two-dimensional data and three-dimensional data of the human body in the target scene, the actions and postures of the personnel in the education scene can be perceived in real time.
[0051] On the other hand, by performing feature fusion processing on the key feature data to obtain a comprehensive feature vector, and performing pose recognition on the comprehensive feature vector based on a preset classification algorithm model to obtain a pose recognition result, the human body pose can be accurately detected through the fusion of two-dimensional data and three-dimensional data. Even in complex backgrounds or multi-person scenarios, the action intentions of each human body can be accurately recognized, improving the overall intelligent level.
[0052] On the other hand, by obtaining the correspondence between the preset pose recognition result and the action intention in the target scene, based on the correspondence and the pose recognition result, the target action intention matching the pose recognition result is obtained. Finally, a control instruction is obtained based on the pose recognition result and the target action intention, and the robot is controlled to execute the corresponding interaction action based on the control instruction. In this way, the present application is not only applicable to traditional classroom teaching, but also can be applied in various scenarios such as extracurricular tutoring, online education, and popular science exhibitions, with wide applicability. And through accurate human body pose detection, the educational robot can perceive the actions and postures of students in real time, execute corresponding interaction actions. The interaction method based on human body pose detection is more natural and intuitive, and users can interact with the robot without complex voice commands or button operations. This interaction method greatly improves the user experience and reduces the usage threshold.
[0053] In an exemplary embodiment provided by the present application, when the robot in the target scene is first subjected to teaching control or it is detected that the surrounding environment has changed, before collecting the two-dimensional data and three-dimensional data, it is necessary to perform initialization processing on the image sensor and the three-dimensional data sensor for collecting human body pose data to make them adapt to the current target scene, thereby improving the accuracy of data recognition processing and control. The specific steps may include:
[0054] Obtain the environmental information of the target scene;
[0055] Based on the environmental information, perform initialization processing on the sensor parameters of the image sensor and the three-dimensional data sensor, obtain the processed sensor parameters and apply them to the image sensor and the three-dimensional data sensor to obtain the human body pose data;
[0056] Obtain the background complexity and the expected interaction method in the target scene;
[0057] Initialize the model parameters of the classification algorithm model and the rule parameters of the preset feature extraction rules respectively based on the background complexity and the expected interaction method, obtain the processed model parameters and the processed rule parameters, and apply them to the classification algorithm model and the preset feature extraction rules.
[0058] Preferably, the image sensor and the 3D data sensor of the present application are preferably a camera and a lidar. In strong light or weak light environments, the parameters of the camera can be dynamically adjusted to ensure image quality; the lidar is not affected by light and can stably detect the 3D contour of the human body. In some application scenarios, the cost of the lidar is relatively high, and its environmental adaptability may be inferior to that of a depth sensor. Therefore, a combination of a camera and a depth sensor (such as an RGB-D camera) can be used to replace the solution of a camera and a lidar. The depth sensor can directly provide 3D information and has a relatively low cost, making it suitable for cost-sensitive educational scenarios. For the convenience of understanding, the embodiments of the present application described in this article use a camera and a lidar as the image sensor and the 3D data sensor for illustration.
[0059] In this embodiment, the device initialization before collecting human body pose data includes the initialization of the image sensor and the 3D data sensor, as well as the initialization of the classification algorithm model and the preset feature extraction rules.
[0060] Specifically, perform initialization processing on the sensor parameters such as the exposure and gain of the camera based on the light conditions included in the environmental information, and perform initialization processing on the sensor parameters such as the scanning frequency, resolution, and detection range of the lidar based on the environmental information, so that the camera and the lidar are adapted to the environment of the current target scene. And adjust the sensitivity parameters of the classification algorithm model and the feature extraction parameters of the preset feature extraction rules according to the background complexity and the expected interaction method of the target scene to optimize the recognition effect.
[0061] In addition, in the embodiments provided by the present application, before collecting 2D data and 3D data, device self-checks can also be performed on the image sensor and the 3D data sensor to improve the reliability of the data. Specifically, check whether the camera as the image sensor can normally collect images, and whether the resolution and frame rate meet the requirements; check whether the lidar as the 3D data sensor can normally emit and receive laser signals, and whether the detection range and accuracy are normal; check the synchronization status between the sensors to ensure that the data of the camera and the lidar can be accurately aligned. If any sensor module fails during the self-check, an alarm will be sent to the user through the robot display screen or voice prompt, indicating the specific fault information, and the subsequent process will be paused until the fault is repaired and the self-check is completed again.
[0062] In an exemplary embodiment provided by the present application, the specific steps of collecting two-dimensional data and three-dimensional data in the target scenario may include:
[0063] Using an image sensor, periodically collect the appearance information and pose information of the human body in the target scenario, and obtain the initial two-dimensional data based on the appearance information and pose information;
[0064] Using a three-dimensional data sensor, periodically collect the point cloud data of the human body in the target scenario to obtain the initial three-dimensional data;
[0065] Perform a first preprocessing operation on the initial two-dimensional data to obtain the processed two-dimensional data. The first preprocessing operation includes denoising processing, contrast enhancement processing, and normalization processing;
[0066] Perform a second preprocessing operation on the initial three-dimensional data to obtain the processed three-dimensional data. The second preprocessing operation includes filtering processing and smoothing processing;
[0067] Determine the human body pose data of the target object according to the processed two-dimensional data and the processed three-dimensional data.
[0068] In this embodiment, control the camera as the image sensor to collect human body image data at a set frame rate, periodically obtain the appearance information and pose information of the human body, and use it as two-dimensional data after the first preprocessing operation. The lidar as the three-dimensional data sensor emits laser signals at a set scanning frequency to obtain the point cloud data of the human body, and use it as three-dimensional data for detecting the contour and spatial position of the human body after the second preprocessing operation.
[0069] In this embodiment, the data preprocessing steps include denoising, contrast enhancement, and normalization processing on the image data collected by the camera as the initial two-dimensional data; filtering and smoothing processing on the point cloud data collected by the lidar as the initial three-dimensional data to remove noise points and outliers to improve the data quality.
[0070] In this way, through the above embodiments, the present application combines the image sensor and the three-dimensional data sensor, gives full play to the advantages of the camera in image recognition and the lidar in three-dimensional space perception, and realizes high-precision detection of human body poses.
[0071] In an exemplary embodiment provided by the present application, before feature extraction of the human body pose data, the two-dimensional data and the three-dimensional data will also be aligned to improve the data quality and the efficiency of subsequent processing, and realize more accurate pose recognition. The specific steps may include:
[0072] Perform spatio-temporal alignment processing on the two-dimensional data and the three-dimensional data to obtain the two-dimensional data and the three-dimensional data with a unified spatio-temporal reference.
[0073] In this embodiment, the data of the camera and the lidar are aligned in space and time to ensure that the data of both can be fused and processed in the same coordinate system. After being processed by the space-time alignment technology in this way, the two-dimensional image data of the camera and the three-dimensional point cloud data of the lidar can be fused and processed. This fusion method can describe the human body posture more comprehensively and provide richer information for subsequent posture recognition.
[0074] In an exemplary embodiment provided by the present application, the specific steps of extracting features from human body posture data based on a preset feature extraction rule may include:
[0075] Based on the preset feature extraction rule, use computer vision algorithms and deep learning models to extract human key point features from two-dimensional data, and obtain human key point features;
[0076] Extract contour features from three-dimensional data to obtain three-dimensional contour features. The three-dimensional contour information includes limb length, limb width, and free position;
[0077] Extract motion features from two-dimensional data and three-dimensional data to obtain motion features. The motion features include posture angle, motion speed, and direction;
[0078] Based on the human key point features, three-dimensional contour features, and motion features, form key feature data.
[0079] In this embodiment, corresponding preset feature extraction rules are configured for different target scenarios. The preset feature extraction rules set the types of key feature data that need to be recognized in the target scenario. The types of key feature data include human key point features, three-dimensional contour features, and motion features. The specific content of the steps for extracting each feature can be, first, determine the feature types to be extracted from two-dimensional data and three-dimensional data based on the preset feature extraction rule, and use computer vision algorithms (such as OpenPose) and deep learning models (such as convolutional neural networks) to extract the positions of human key points (such as head, shoulders, hands, legs, etc.) from the image data as two-dimensional data to obtain human key point features; extract the three-dimensional contour features of the human body from the point cloud data collected by the lidar as three-dimensional data, including the length, width, and spatial position of the limbs; combine the image data and point cloud data of the camera and the lidar to extract the posture angle, motion speed, and direction of the human body, etc., as motion features.
[0080] In another exemplary embodiment provided by the present application, the specific steps of performing feature fusion processing on the key feature data to obtain a comprehensive feature vector may include:
[0081] Obtain the environmental information of the target scenario, the first reliability parameter corresponding to the image sensor that collects two-dimensional data, and the second reliability parameter corresponding to the three-dimensional data sensor that collects three-dimensional data;
[0082] Based on the environmental information, the first reliability parameter, and the second reliability parameter, determine the respective fusion weights corresponding to the image sensor and the three-dimensional data sensor;
[0083] Based on the fusion weights, perform weighted fusion on the human key point features and the three-dimensional contour features to obtain weighted fusion features;
[0084] Through a preset data fusion algorithm, perform fusion processing on the motion features to obtain multi-modal fusion features;
[0085] Based on the weighted fusion features and the multi-modal fusion features, form a comprehensive feature vector.
[0086] In this embodiment, the comprehensive feature vector includes the weighted fusion features corresponding to the human key point features and the three-dimensional contour features, and the multi-modal fusion features corresponding to the motion features.
[0087] Specifically, the human key point features and the three-dimensional contour features are weighted and fused according to the preset fusion weights to obtain weighted fusion features, and the fusion weights are dynamically adjusted according to the reliability parameters of the sensors and the environmental information. Since the motion features are obtained by extracting features from two-dimensional data and three-dimensional data, they are multi-modal features. Multi-modal features refer to a set of features that are extracted from different types of data modalities and can comprehensively represent information. In this embodiment, the Kalman filter or other data fusion algorithms are used to perform fusion processing on the multi-modal features to improve the accuracy and stability of the features and provide reliable input for subsequent pose recognition.
[0088] In an exemplary embodiment provided by the present application, the specific steps for performing pose recognition on the comprehensive feature vector based on a preset classification algorithm model may include:
[0089] Perform pose recognition on the comprehensive feature vector based on a preset classification algorithm model to obtain pose actions, and the classification algorithm model is a support vector machine or a deep neural network;
[0090] Based on the time tags corresponding to the pose actions, sort the pose actions in chronological order to obtain a sorting result;
[0091] Based on the sorting result and the pose actions, identify continuous actions;
[0092] Based on the pose actions and the continuous actions, form a pose recognition result.
[0093] In this embodiment, a support vector machine (SVM), a deep neural network (DNN), or other classification algorithms are used to perform pose recognition on the comprehensive feature vector to obtain pose actions, so as to determine the pose type of the human body, such as standing, sitting, raising a hand, waving, turning around, etc. For complex poses, the recognized human poses are sorted based on the corresponding time tags to obtain a sorting result, and continuous actions (such as the start and end of waving) are recognized in combination with the time series.
[0094] In this way, through the above embodiments, the present application improves the accuracy of pose recognition by recognizing individual pose actions and continuous actions formed by multiple pose actions, thereby improving the accuracy of controlling the robot and better realizing interaction.
[0095] In an exemplary embodiment provided by the present application, the pose recognition result includes at least one unit action, and the unit action is a pose action or a continuous action. The specific steps for obtaining a control instruction based on the pose recognition result and the target action intention may include:
[0096] If the pose recognition result includes only one unit action, a control instruction is obtained based on the target action intention corresponding to the unit action;
[0097] If the pose recognition result includes at least two unit actions, the association information between every two adjacent unit actions in time is obtained;
[0098] A control instruction is obtained based on the association information and the target action intention corresponding to each unit action.
[0099] In this embodiment, when the pose recognition result includes at least two unit actions, the association information between every two adjacent unit actions in time is obtained, and a control instruction is obtained based on the association information and the target action intention corresponding to each unit action, so as to realize the smooth switching of two interaction actions of the robot and more accurate interaction.
[0100] In this application, the correspondence between the preset posture recognition results and action intentions in the target scenario may include: when the "left hand lifted" posture is recognized, it is determined to control the robot to turn left; when the "left hand continuously waved" posture is recognized, it is determined to control the robot to move horizontally to the left; when the "right hand lifted" posture is recognized, it is determined to control the robot to turn right; when the "right hand continuously waved" posture is recognized, it is determined to control the robot to move horizontally to the right; when the "hands crossed in front of the chest" posture is recognized, it is determined to calibrate the target and only recognize the current target posture among multiple targets; when the "palms of both hands forward" posture is recognized, it is determined to control the robot to move forward; when the "one hand forward" posture is recognized, it is determined to pause the movement of the device; when the "hands moving back and forth" posture is recognized, it is determined to control the robot to move backward.
[0101] In an exemplary embodiment provided by this application, after controlling the robot to execute the corresponding action based on the control instruction, the execution state of the robot can be continuously monitored and optimized based on this. The specific steps may include:
[0102] Obtain the feedback information and execution result generated by the robot based on the control instruction;
[0103] Update the preset feature extraction rule based on the feedback information and execution result to obtain the updated feature extraction rule, and apply it to the feature extraction of new two-dimensional data and three-dimensional data.
[0104] In this embodiment, when it is determined that the robot has made a misjudgment or response delay based on the feedback information and execution result, the preset feature extraction rule is updated based on the feedback information and execution result to obtain the updated feature extraction rule, which is applied to the feature extraction of new two-dimensional data and three-dimensional data, and this error is recorded and sent to the user. According to the user's feedback and actual usage situation, the generation logic of the control instruction is optimized to improve the robustness of the system and the user experience.
[0105] In this way, through the above embodiments, this application can dynamically adjust the posture recognition algorithm and control strategy according to the real-time monitoring results, and continuously optimize its own performance. This adaptive ability enables the robot to better adapt to different teaching environments and user needs during long-term use.
[0106] This enables this application to not only adapt to different target scenarios and environmental conditions, but also continuously improve its own performance through dynamic optimization, providing a more intelligent and natural teaching tool for the education field.
[0107] Figure 2 It is a block diagram of a robot control system 200 based on human posture detection shown in an exemplary embodiment of this application. As Figure 2 shown, this system includes:
[0108] The acquisition unit 201 is used to obtain the human body pose data of the target object in the target scene, and the human body pose data includes two-dimensional data and three-dimensional data;
[0109] The feature extraction unit 202 is used to extract features from the human body pose data based on a preset feature extraction rule to obtain the key feature data corresponding to the target object, and the key feature data includes at least one type of human body pose feature corresponding to the target object;
[0110] The data fusion unit 203 is used to perform feature fusion processing on the key feature data to obtain a comprehensive feature vector;
[0111] The recognition unit 204 is used to perform pose recognition on the comprehensive feature vector based on a preset classification algorithm model to obtain a pose recognition result;
[0112] The matching unit 205 is used to obtain the correspondence between the preset pose recognition result and the action intention in the target scene, and based on the correspondence and the pose recognition result, obtain the target action intention matching the pose recognition result;
[0113] The processing unit 206 is used to obtain a control instruction based on the pose recognition result and the target action intention, so as to control the robot to execute corresponding interaction actions based on the control instruction.
[0114] This system applies the robot control method based on human body pose detection provided by this application. The acquisition unit 201 acquires the two-dimensional data and three-dimensional data of the human body in the target scene. The feature extraction unit 202 extracts features from the two-dimensional data and three-dimensional data based on a preset feature extraction rule to obtain the key feature data corresponding to the preset feature extraction rule, and the key feature data includes at least one type of human body pose feature corresponding to the target object. Then, the data fusion unit 203 performs feature fusion processing on the key feature data to obtain a comprehensive feature vector; the recognition unit 204 performs pose recognition on the comprehensive feature vector based on a preset classification algorithm model to obtain a pose recognition result; the matching unit 205 obtains the correspondence between the preset pose recognition result and the action intention in the target scene, and based on the correspondence and the pose recognition result, obtains the target action intention matching the pose recognition result. Finally, the processing unit 206 obtains a control instruction based on the pose recognition result and the target action intention, so as to control the robot to execute corresponding interaction actions based on the control instruction.
[0115] In this way, this application can real-time sense the actions and postures of people in the education scene. Through the fusion of two-dimensional data and three-dimensional data, while achieving accurate human body pose detection, it reduces the influence of the environment on pose recognition to adapt to different target scenes and environmental conditions, and according to the correspondence between the preset pose recognition result and the action intention in the target scene, provides personalized teaching content by controlling the robot.
[0116] In another exemplary embodiment, the system further includes:
[0117] An initialization unit, configured to obtain environmental information of a target scenario; initialize sensor parameters of an image sensor and a three-dimensional data sensor based on the environmental information, obtain processed sensor parameters and apply them to the image sensor and the three-dimensional data sensor to obtain human body pose data; obtain background complexity and an expected interaction mode in the target scenario; initialize model parameters of a classification algorithm model and rule parameters of a preset feature extraction rule respectively based on the background complexity and the expected interaction mode, obtain processed model parameters and processed rule parameters, and apply them to the classification algorithm model and the preset feature extraction rule.
[0118] In another exemplary embodiment, the acquisition unit 201 is further configured to use the image sensor to periodically acquire appearance information and pose information of a human body in the target scenario, and obtain initial two-dimensional data based on the appearance information and the pose information; use the three-dimensional data sensor to periodically acquire point cloud data of the human body in the target scenario to obtain initial three-dimensional data; perform a first preprocessing operation on the initial two-dimensional data to obtain processed two-dimensional data, where the first preprocessing operation includes denoising processing, contrast enhancement processing, and normalization processing; perform a second preprocessing operation on the initial three-dimensional data to obtain processed three-dimensional data, where the second preprocessing operation includes filtering processing and smoothing processing; determine human body pose data of a target object according to the processed two-dimensional data and the processed three-dimensional data.
[0119] In another exemplary embodiment, the system further includes:
[0120] An alignment unit, configured to perform spatio-temporal alignment processing on the two-dimensional data and the three-dimensional data to obtain two-dimensional data and three-dimensional data with a unified spatio-temporal reference.
[0121] In another exemplary embodiment, the feature extraction unit 202 is further configured to, based on a preset feature extraction rule, use a computer vision algorithm and a deep learning model to extract human body key point features from the two-dimensional data to obtain human body key point features; extract contour features from the three-dimensional data to obtain three-dimensional contour features, where the three-dimensional contour features include limb length, limb width, and free positions; extract motion features from the two-dimensional data and the three-dimensional data to obtain motion features, where the motion features include pose angles, motion speeds, and directions; form key feature data based on the human body key point features, the three-dimensional contour features, and the motion features.
[0122] In another exemplary embodiment, the data fusion unit 203 is further configured to obtain the environmental information of the target scenario, the first reliability parameter corresponding to the image sensor that collects two-dimensional data, and the second reliability parameter corresponding to the three-dimensional data sensor that collects three-dimensional data; determine the respective fusion weights of the image sensor and the three-dimensional data sensor based on the environmental information, the first reliability parameter, and the second reliability parameter; perform weighted fusion on the human key point features and the three-dimensional contour features based on the fusion weights to obtain weighted fusion features; perform fusion processing on the motion features through a preset data fusion algorithm to obtain multi-modal fusion features; and form a comprehensive feature vector based on the weighted fusion features and the multi-modal fusion features.
[0123] In another exemplary embodiment, the recognition unit 204 is further configured to perform pose recognition on the comprehensive feature vector based on a preset classification algorithm model to obtain pose actions, where the classification algorithm model is a support vector machine or a deep neural network; sort the pose actions in chronological order based on the time tags corresponding to the pose actions to obtain a sorting result; recognize continuous actions based on the sorting result and the pose actions; and form a pose recognition result based on the pose actions and the continuous actions.
[0124] In another exemplary embodiment, the pose recognition result includes at least one unit action, and the unit action is a pose action or a continuous action; the processing unit 206 is further configured to, if the pose recognition result includes only one unit action, obtain a control instruction based on the target action intention corresponding to the unit action; if the pose recognition result includes at least two unit actions, obtain the association information between every two adjacent unit actions in time; and obtain a control instruction based on the association information and the target action intention corresponding to each unit action.
[0125] In another exemplary embodiment, the system further includes:
[0126] An update unit, configured to obtain the feedback information and execution result generated by the robot based on the control instruction; update a preset feature extraction rule based on the feedback information and the execution result to obtain an updated feature extraction rule, and apply it to the feature extraction of new two-dimensional data and three-dimensional data.
[0127] It should be noted that the robot control system based on human pose detection provided in the above embodiment and the robot control method based on human pose detection provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the robot control system based on human pose detection provided in the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above, and this will not be limited here either.
[0128] Embodiments of the present application also provide an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the robot control method based on human pose detection provided in each of the above embodiments.
[0129] Figure 3 The structural schematic diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. It should be noted that, Figure 3 The computer system 300 of the electronic device shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present application.
[0130] As Figure 3 shown, the computer system 300 includes a central processing unit (CPU) 301, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 302 or the program loaded from the storage section 308 into the random access memory (RAM) 303, such as executing the method in the above embodiments. In the RAM 303, various programs and data required for system operation are also stored. The CPU 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0131] The following components are connected to the I / O interface 305: an input section 306 including a keyboard, a mouse, etc.; an output section 307 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 308 including a hard disk, etc.; and a communication section 309 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 309 performs communication processing via a network such as the Internet. A drive 310 is also connected to the I / O interface 305 as required. A removable medium 311, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 310 as required, so that the computer program read from it can be installed into the storage section 308 as required.
[0132] In particular, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 309, and / or installed from the removable medium 311. When the computer program is executed by the central processing unit (CPU) 301, various functions defined in the system of the present application are executed.
[0133] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can, for example, be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program included on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.
[0134] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0135] The units involved in the embodiments described in the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the units themselves in some cases.
[0136] On the other hand, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the robot control method based on human pose detection as described above. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device.
[0137] On the other hand, the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the robot control method based on human pose detection provided in the above various embodiments.
[0138] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, or improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A robot control method based on human body posture detection, characterized in that, The method includes: Obtaining human body pose data of a target object in a target scenario, where the human body pose data includes two-dimensional data and three-dimensional data; Performing feature extraction on the human body pose data based on a preset feature extraction rule to obtain key feature data corresponding to the target object, where the key feature data includes at least one type of human body pose feature corresponding to the target object; Performing feature fusion processing on the key feature data to obtain a comprehensive feature vector; Performing pose recognition on the comprehensive feature vector based on a preset classification algorithm model to obtain a pose recognition result; Obtaining the correspondence between a preset pose recognition result and an action intention in the target scenario, and based on the correspondence and the pose recognition result, obtaining a target action intention that matches the pose recognition result; Obtaining a control instruction based on the pose recognition result and the target action intention, and controlling a robot to execute a corresponding interaction action based on the control instruction.
2. The method according to claim 1, wherein Before obtaining the human body pose data of the target object in the target scenario, the method further includes: Obtaining environmental information of the target scenario; Performing initialization processing on sensor parameters of an image sensor and a three-dimensional data sensor based on the environmental information, obtaining the processed sensor parameters and applying them to the image sensor and the three-dimensional data sensor to obtain the human body pose data; Obtaining the background complexity and an expected interaction mode in the target scenario; Performing initialization processing on model parameters of the classification algorithm model and rule parameters of the preset feature extraction rule respectively based on the background complexity and the expected interaction mode, obtaining the processed model parameters and the processed rule parameters, and applying them to the classification algorithm model and the preset feature extraction rule.
3. The method according to claim 1, wherein The obtaining of the human body pose data of the target object in the target scenario, where the human body pose data includes two-dimensional data and three-dimensional data, includes: Using an image sensor to periodically collect appearance information and pose information of a human body in the target scenario, and obtaining initial two-dimensional data based on the appearance information and the pose information; Using a three-dimensional data sensor to periodically collect point cloud data of the human body in the target scenario to obtain initial three-dimensional data; Performing a first preprocessing operation on the initial two-dimensional data to obtain processed two-dimensional data, where the first preprocessing operation includes denoising processing, contrast enhancement processing, and normalization processing; Performing a second preprocessing operation on the initial three-dimensional data to obtain processed three-dimensional data, where the second preprocessing operation includes filtering processing and smoothing processing; Determining the human body pose data of the target object according to the processed two-dimensional data and the processed three-dimensional data.
4. The method according to claim 1, wherein Before performing feature extraction on the human body pose data based on a preset feature extraction rule to obtain the key feature data corresponding to the target object, the method further includes: Performing spatio-temporal alignment processing on the two-dimensional data and the three-dimensional data to obtain two-dimensional data and three-dimensional data with a unified spatio-temporal reference.
5. The method according to claim 1, wherein The performing of feature extraction on the human body pose data based on a preset feature extraction rule to obtain key feature data includes: Based on the preset feature extraction rules, use computer vision algorithms and deep learning models to extract human key-point features from the two-dimensional data, obtaining human key-point features; Extract contour features from the three-dimensional data, obtaining three-dimensional contour features, where the three-dimensional contour features include limb length, limb width, and free positions; Extract motion features from the two-dimensional data and the three-dimensional data, obtaining motion features, where the motion features include pose angles, motion speeds, and directions; Based on the human key-point features, the three-dimensional contour features, and the motion features, form key feature data.
6. The method according to claim 5, wherein Perform feature fusion processing on the key feature data to obtain a comprehensive feature vector, including: Obtain the environmental information of the target scene, the first reliability parameter corresponding to the image sensor that collects the two-dimensional data, and the second reliability parameter corresponding to the three-dimensional data sensor that collects the three-dimensional data; Based on the environmental information, the first reliability parameter, and the second reliability parameter, determine the respective fusion weights corresponding to the image sensor and the three-dimensional data sensor; Based on the fusion weights, perform weighted fusion on the human key-point features and the three-dimensional contour features, obtaining a weighted fusion feature; Through a preset data fusion algorithm, perform fusion processing on the motion features, obtaining a multi-modal fusion feature; Based on the weighted fusion feature and the multi-modal fusion feature, form a comprehensive feature vector.
7. The method according to claim 1, wherein Perform pose recognition on the comprehensive feature vector based on a preset classification algorithm model, obtaining a pose recognition result, including: Perform pose recognition on the comprehensive feature vector based on a preset classification algorithm model, obtaining pose actions, where the classification algorithm model is a support vector machine or a deep neural network; Based on the time tags corresponding to the pose actions, sort the pose actions in chronological order, obtaining a sorting result; Based on the sorting result and the pose actions, identify continuous actions; Based on the pose actions and the continuous actions, form a pose recognition result.
8. The method according to claim 1, wherein The pose recognition result includes at least one unit action, where the unit action is a pose action or a continuous action; obtaining a control instruction based on the pose recognition result and the target action intention includes: If the pose recognition result includes only one unit action, obtain a control instruction based on the target action intention corresponding to the unit action; If the pose recognition result includes at least two unit actions, obtain the association information between every two adjacent unit actions in time; Based on the association information and the target action intentions corresponding to the unit actions respectively, obtain a control instruction.
9. The method according to claim 1, wherein The method further includes: Obtain the feedback information and execution result generated by the robot based on the control instruction; Based on the feedback information and the execution result, update the preset feature extraction rules, obtaining updated feature extraction rules, and apply them to the feature extraction of new two-dimensional data and three-dimensional data.
10. A robot control system based on human pose detection, characterized in that, Including: An acquisition unit, configured to acquire the human body pose data of a target object in a target scene, where the human body pose data includes two-dimensional data and three-dimensional data; A feature extraction unit, configured to perform feature extraction on the human body pose data based on a preset feature extraction rule, so as to obtain key feature data corresponding to the target object, where the key feature data includes at least one type of human body pose feature corresponding to the target object; A data fusion unit, configured to perform feature fusion processing on the key feature data to obtain a comprehensive feature vector; An identification unit, configured to perform pose identification on the comprehensive feature vector based on a preset classification algorithm model to obtain a pose identification result; A matching unit, configured to obtain the correspondence between the preset pose identification result and the action intention in the target scenario, and based on the correspondence and the pose identification result, obtain a target action intention matching the pose identification result; A processing unit, configured to obtain a control instruction based on the pose identification result and the target action intention, so as to control the robot to execute corresponding interaction actions based on the control instruction.