Posture Recognition Method, Device, System, Equipment, Medium, Product and Rehabilitation Mirror

By combining infrared images and depth image data, a three-dimensional skeleton sequence diagram is generated using neural network models, which solves the limitations of posture recognition technology in terms of direction and angle, and achieves high accuracy and stability of posture recognition, which is suitable for real-time monitoring and evaluation in rehabilitation training.

CN119206785BActive Publication Date: 2025-08-01STAR SPORTS MEDICINE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202411324478.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-23
Publication Date
2025-08-01
Estimated Expiration
2044-09-23

AI Technical Summary

Technical Problem

The existing pose recognition technology has limitations in recognition direction and angle, the recognition process is complex and the accuracy is not high, and the pose data is unstable.

Method used

By combining infrared image data and depth image data, a pre-trained object structure recognition model is used to generate a three-dimensional skeleton sequence diagram, and a neural network model is used to fit the data to identify postures in any direction and angle.

Benefits of technology

Improves the accuracy and stability of posture recognition, and can recognize postures in any direction and angle, which is suitable for real-time monitoring and evaluation in rehabilitation training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119206785B_ABST
    Figure CN119206785B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention discloses a posture recognition method, device, system, equipment, medium, product, and rehabilitation mirror. Among them, the method includes: obtaining infrared image data and depth image data of a target object, where the infrared image data and the depth image data are a pair of sequence image data synchronously collected within a preset duration; inputting the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtaining an object structure segmentation map based on the object structure recognition result; performing data fitting on the object structure segmentation map and the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object. The technical solution of the embodiment of the present invention solves the problems of inaccurate current posture recognition and limitations in recognition direction and angle. By combining two-dimensional and three-dimensional image data and a neural network model, a three-dimensional skeleton sequence map can be generated, which can recognize postures in any direction and angle, improving the accuracy and stability of posture recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the technical field of computer vision, and in particular, to a pose recognition method, device, system, equipment, medium, product, and rehabilitation mirror. Background Art

[0002] With the development of computer technology, pose recognition has played an important role in many fields such as motion analysis and sports, medical and rehabilitation, and entertainment.

[0003] However, the current pose recognition has limitations. It can only recognize in its specified directions and angles. The recognition process is complex and the recognition accuracy is not high. The obtained pose data is also more unstable. Summary of the Invention

[0004] Embodiments of the present invention provide a pose recognition method, device, system, equipment, medium, product, and rehabilitation mirror. By combining two-dimensional and three-dimensional image data and a neural network model, a three-dimensional skeleton sequence diagram can be generated, which can recognize poses in any direction and angle, improving the accuracy and stability of pose recognition.

[0005] In a first aspect, embodiments of the present invention provide a pose recognition method, which includes:

[0006] Obtain infrared image data and depth image data of a target object, where the infrared image data and the depth image data are a pair of sequence image data synchronously collected within a preset time period;

[0007] Input the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtain an object structure segmentation map based on the object structure recognition result;

[0008] Perform data fitting on the object structure segmentation map and the corresponding depth image data to obtain a three-dimensional skeleton sequence diagram of the target object within the preset time period.

[0009] In a second aspect, embodiments of the present invention provide a pose recognition device, which includes:

[0010] An image acquisition module, configured to obtain infrared image data and depth image data of a target object, where the infrared image data and the depth image data are a pair of sequence image data synchronously collected within a preset time period;

[0011] An object structure segmentation map generation module, configured to input the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtain an object structure segmentation map based on the object structure recognition result;

[0012] The 3D skeleton sequence diagram generation module is used to perform data fitting between the object structure segmentation diagram and the corresponding depth image data to obtain a 3D skeleton sequence diagram of the target object within a preset time length.

[0013] In a third aspect, an embodiment of the present invention provides a gesture recognition system, the system comprising:

[0014] 3D depth camera, infrared camera and computing processing platform;

[0015] Wherein, the three-dimensional depth camera is used to collect depth image data of the target object within a preset time;

[0016] An infrared camera, used to collect infrared image data of a target object within a preset time;

[0017] The computing and processing platform is used to perform gesture recognition based on the depth image data and the infrared image data according to the method provided in any embodiment of the present invention.

[0018] In a fourth aspect, an embodiment of the present invention further provides a computer device, comprising:

[0019] one or more processors;

[0020] a memory for storing one or more programs;

[0021] When the one or more programs are executed by one or more processors, the one or more processors implement the gesture recognition method provided by any embodiment of the present invention.

[0022] In a fifth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the gesture recognition method provided by any embodiment of the present invention.

[0023] In a sixth aspect, an embodiment of the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the gesture recognition method provided by any embodiment of the present invention.

[0024] In a seventh aspect, an embodiment of the present invention further provides a rehabilitation mirror, characterized in that it includes:

[0025] Image acquisition module, calculation processing module and display screen;

[0026] The image acquisition module is used to acquire depth image data and infrared image data of the target object within a preset time;

[0027] The calculation and processing module is used to perform pose recognition based on depth image data and infrared image data according to the pose recognition method provided in any embodiment of the present invention, and generate a three-dimensional skeleton sequence diagram and / or a three-dimensional skeleton diagram;

[0028] The display screen is used to display the three-dimensional skeleton sequence diagram and / or the three-dimensional skeleton diagram.

[0029] The embodiments in the above invention have the following advantages or beneficial effects:

[0030] In the embodiment of the present invention, by acquiring infrared image data and depth image data of a target object, wherein the infrared image data and the depth image data are a pair of sequence image data synchronously acquired within a preset time period; inputting the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtaining an object structure segmentation diagram based on the object structure recognition result; performing data fitting on the object structure segmentation diagram and the corresponding depth image data to obtain a three-dimensional skeleton sequence diagram of the target object within the preset time period. The technical solution of the embodiment of the present invention solves the problems of inaccurate current pose recognition and limitations in recognition direction and angle. By combining two-dimensional and three-dimensional image data and a neural network model, a three-dimensional skeleton sequence diagram can be generated, and poses in any direction and angle can be recognized, improving the accuracy and stability of pose recognition. Description of the Drawings

[0031] Figure 1 is a flowchart of a pose recognition method provided by an embodiment of the present invention;

[0032] Figure 2 is a flowchart of a pose recognition method provided by an embodiment of the present invention;

[0033] Figure 3 is a schematic diagram of joint space information storage provided by an embodiment of the present invention;

[0034] Figure 4 is a flowchart of a pose recognition method provided by an embodiment of the present invention;

[0035] Figure 5 is a schematic diagram of target object determination provided by an embodiment of the present invention;

[0036] Figure 6 is a schematic structural diagram of a pose recognition device provided by an embodiment of the present invention;

[0037] Figure 7 is a schematic structural diagram of a pose recognition system provided by an embodiment of the present invention;

[0038] Figure 8 is a schematic functional diagram of a pose recognition system provided by an embodiment of the present invention;

[0039] Figure 9 It is a schematic structural diagram of a rehabilitation mirror provided by an embodiment of the present invention;

[0040] Figure 10 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Detailed implementation manners

[0041] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only parts related to the present invention rather than all structures are shown in the accompanying drawings.

[0042] Figure 1 It is a flowchart of a posture recognition method provided by an embodiment of the present invention, and this embodiment is applicable to the scenario of posture recognition. This method can be executed by a posture recognition device, and this device can be implemented in a software and / or hardware manner and integrated in a computer device with application development functions.

[0043] As Figure 1 shown, the posture recognition method of this embodiment includes the following steps:

[0044] S110. Obtain infrared image data and depth image data of a target object.

[0045] Among them, the infrared image data and the depth image data are a pair of sequence image data synchronously collected within a preset time period.

[0046] The target object can be an object that needs rehabilitation training or whose body posture needs to be recognized, such as a human body or other livestock and pets, etc.

[0047] Obtain a pair of sequence image data of the target object synchronously collected by an infrared camera and a depth camera within a preset time period. The preset time period can be set according to actual needs and can be determined by the actual acquisition time. This embodiment does not make any limitations in this regard. A pair of sequence image data includes an infrared image and a depth image corresponding to the acquisition moment.

[0048] In this embodiment, the depth camera can be composed of a three-dimensional depth sensor based on the time-of-flight ranging principle and a color camera, and has good depth image data acquisition capabilities.

[0049] S120. Input the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtain an object structure segmentation map based on the object structure recognition result.

[0050] The pre-trained object structure recognition model can be a deep learning-based neural network model. Through training and optimization with a large amount of data, it has learned the features of the target object joints and various body parts in the object structure. When infrared image data is input, the object structure recognition model can quickly identify each joint of the target object, such as the elbow joint, knee joint, and ankle joint, etc. At the same time, the object structure recognition model can also perform fine segmentation on each part of the target object, dividing the target into different regions such as the head, neck, torso, arms, hands, legs, and feet. According to the joint recognition result and the part recognition result, a two-dimensional segmentation map of the object joints and each part of the object is obtained.

[0051] S130. Fit the object structure segmentation map with the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object within a preset time period.

[0052] Use the depth image data to fit with the two-dimensional object joints and the segmentation map of each part of the object, so as to obtain object joints with three-dimensional coordinates and form a skeleton, and obtain a three-dimensional skeleton sequence map corresponding to each acquisition moment within a preset time period of the target object. The three-dimensional skeleton sequence map can reflect the body posture of the target object at each corresponding acquisition moment.

[0053] The technical solution of this embodiment, by obtaining the infrared image data and depth image data of the target object, wherein the infrared image data and depth image data are a pair of sequence image data synchronously acquired within a preset time period; inputting the infrared image data into the pre-trained object structure recognition model to obtain an object structure recognition result, and obtaining an object structure segmentation map based on the object structure recognition result; fitting the object structure segmentation map with the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object within a preset time period. The technical solution of the embodiment of the present invention solves the problems of inaccurate current posture recognition and limitations in recognition direction and angle. It can generate a three-dimensional skeleton sequence map by combining two-dimensional and three-dimensional image data and a neural network model, and can recognize postures in any direction and angle, improving the accuracy and stability of posture recognition.

[0054] Figure 2 It is a flowchart of a posture recognition method provided by an embodiment of the present invention. This embodiment and the posture recognition method in the above embodiment belong to the same inventive concept, and further describes the process of updating the three-dimensional skeleton map. This method can be executed by a posture recognition device, and the device can be implemented in a software and / or hardware manner and integrated in a computer device with application development functions.

[0055] As Figure 2 shown, the posture recognition method of this embodiment includes the following steps:

[0056] S210. Obtain the infrared image data and depth image data of the target object.

[0057] Among them, the infrared image data and depth image data are a pair of sequence image data synchronously collected within a preset time period.

[0058] S220. Input the infrared image data into the pre-trained object structure recognition model to obtain the object structure recognition result, and obtain the object structure segmentation map based on the object structure recognition result.

[0059] S230. Perform data fitting on the object structure segmentation map and the corresponding depth image data to obtain the three-dimensional skeleton sequence map of the target object within the preset time period.

[0060] S240. Obtain the first coordinate observation value of each joint in each three-dimensional skeleton map in the three-dimensional skeleton sequence map in the camera coordinate system.

[0061] Read the first coordinate observation value of each joint in each three-dimensional skeleton map in the three-dimensional skeleton sequence map through the coordinate reading function and store it in the preset storage location. The first coordinate observation value can include the first position coordinate observation value and the first direction coordinate observation value.

[0062] The spatial position coordinates of specific object joints, that is, the first position coordinate observation value, and the joint direction represented by quaternion, that is, the first direction coordinate observation value, can be obtained from the preset storage location according to the identifier of the object. The joints form a skeleton according to the hierarchical structure and connection relationship. Taking human body tracking and skeleton recognition as an example, the skeleton includes a total of 32 joints. The joint hierarchical structure is distributed from the root joint at the center of the human body to the limbs. Each connection (bone) links the parent joint and the child joint. Its joint connection and spatial information format are shown in Table 1 below. Table 1 is the joint connection and joint spatial information table. Among them, the first three elements x, y, z in the joint spatial information [x, y, z, w, x, y, z] represent the spatial position coordinates of the joint, that is, the first position coordinate observation value, and the last four elements w, x, y, z are the joint directions represented by the normalized quaternion, that is, the first direction coordinate observation value.

[0063] Table 1

[0064]

[0065]

[0066]

[0067] S250. Convert the first coordinate observation value to obtain the second coordinate observation value of each joint in the object coordinate system.

[0068] Among them, the second coordinate observation value includes the second position coordinate observation value.

[0069] The machine vision technology for rehabilitation training monitoring is mainly based on pose recognition technology. The machine vision technology can evaluate the patient's motor ability by analyzing the patient's movement trajectory and pose, and provide a personalized exercise training plan by real-time monitoring the patient's movement state. Whether it is the evaluation of the patient's motor ability based on machine vision or the monitoring of the patient's movement state, it is patient-centered. However, the origin of the original coordinate system of a general three-dimensional depth camera is located at the camera focus, and the human joint and skeleton coordinates in the original camera coordinate system are not suitable for functions such as rehabilitation training monitoring. Therefore, in this embodiment, the original camera coordinate system is converted to the object coordinate system of the user or patient. Specifically, the coordinate of a specific joint of the object can be selected as the reference coordinate point, and the camera coordinate system with the coordinate origin at the camera focus is converted to the second coordinate observation value in the object coordinate system with the reference coordinate point as the origin.

[0070] In an alternative embodiment, the first coordinate observation value is converted to obtain the second coordinate observation value of each joint in the object coordinate system, which may be to determine the relative position relationship between the coordinate reference joint in each three-dimensional skeleton diagram and the origin in the object coordinate system; according to the first coordinate observation value of the coordinate reference joint in the camera coordinate system and the relative position relationship, determine the coordinate system mapping relationship between the camera coordinate system and the object coordinate system; according to the coordinate system mapping relationship and the position relationship between the non-coordinate reference joint and the corresponding coordinate reference joint in each three-dimensional skeleton diagram, convert the first coordinate observation value of the non-coordinate reference joint to obtain the second coordinate observation value of each joint in each three-dimensional skeleton diagram in the object coordinate system.

[0071] The coordinate reference joint may be a joint located at the center position of the object skeleton. According to the first coordinate observation value of the coordinate reference joint in the camera coordinate system and the relative position relationship with the origin in the object coordinate system, the coordinate system mapping relationship between the camera coordinate system and the object coordinate system is determined by algorithms such as matrix transformation.

[0072] Taking the target object as a human body as an example, the coordinate reference joint may be the pelvis joint. Let the coordinate of any joint other than the root joint, that is, the pelvis joint, of the human body in the original camera coordinate system be t camera [x, y, z], and the direction be q camera [w, x, y, z]; the position coordinate in the second coordinate observation value of this joint in the human coordinate system is t body [x, y, z], and the direction coordinate is q body [w, x, y, z]. Then the conversion method of converting the coordinate of this joint in the original camera coordinate system to the human coordinate system is:

[0073] t body= t camera -t;

[0074] q body = q × q camera 。

[0075] Wherein, q × q camera is the multiplication of quaternions of direction coordinates.

[0076] It can be understood that although the pose recognition in the three-dimensional space coordinate system can obtain the object joints with three-dimensional coordinates and form a skeleton, there are joint tracking errors caused by factors such as depth information noise, occlusion, and low confidence, resulting in the spatial information of the joints and the skeleton being not reasonable and accurate enough. Therefore, it is necessary to filter the spatial information of the joints, that is, the first coordinate observation value of the joints in the camera coordinate system.

[0077] Therefore, in an alternative embodiment, before converting the first coordinate observation value, the first coordinate observation value is compared and analyzed with a pre-determined sample coordinate observation value to obtain a comparison and analysis result; the outliers in the first coordinate observation value are determined according to the comparison and analysis result and removed; wherein, the sample coordinate observation value is obtained by synchronously collecting infrared image data and depth image data of the sample object in a static and moving state within a preset sample coordinate acquisition duration and reading the coordinates.

[0078] Since the random systematic error follows a Gaussian distribution, it is possible to experimentally obtain the sample data of the spatial position information of the human joint data before recognizing the pose and fit it to the Gaussian distribution. For example, the human body or the human model is placed in the effective working space of the camera and kept stationary, and the system cyclically reads the position coordinates of any joint in the human coordinate system within a period of time and saves them to the data sequence of this joint. The human body is moved multiple times in the effective working space of the camera and the spatial position coordinates of the same joint are repeatedly read and added to the aforementioned data sequence. This data sequence sample generally conforms to the Gaussian distribution, and the Gaussian distribution: N(μ,σ) is fitted, and the mean value μ and the standard deviation σ of this distribution are saved.

[0079] When reading the spatial position coordinate information of the joint, first compare whether the deviation of the coordinate value from the mean value μ of the above Gaussian distribution exceeds 3 times the standard deviation σ, and the data points exceeding 3 times the standard deviation are filtered out as outliers.

[0080] S260. Based on the second coordinate observation value, perform coordinate prediction to obtain the predicted value of the second position coordinate of the corresponding joint in the adjacent next three-dimensional skeleton diagram.

[0081] Such as Figure 3As shown, since the model calculates the predicted value of the second position coordinate of the corresponding joint in the adjacent next three-dimensional skeleton image, the joint space information of the current three-dimensional skeleton image is required, so all joint space information of the human body is saved starting from the initial time t0; this application uses a queue data structure to save all joint space information of the human body in the current three-dimensional skeleton image corresponding to time t-1 and the adjacent next three-dimensional skeleton image corresponding to time t.

[0082] Based on the second coordinate observation, the Kalman filter algorithm performs coordinate prediction to obtain the second coordinate prediction value of the corresponding joint in the next adjacent 3D skeleton graph. The Kalman filter algorithm can calculate the second position coordinate prediction value of the corresponding joint in the next adjacent 3D skeleton graph, i.e., the 3D skeleton graph at time t, by establishing the system observation equation and state transition equation based on the second coordinate observation value of the joint in the 3D skeleton graph at time (t-1).

[0083] In an optional embodiment, coordinate prediction is performed based on the second coordinate observation value to obtain the second position coordinate prediction value of the corresponding joint in the adjacent next three-dimensional skeleton graph. The joint connection vector between each joint and the corresponding parent joint can be determined based on the second coordinate observation value of each joint and the corresponding parent joint; the second position coordinate prediction value of the corresponding joint in the adjacent next three-dimensional skeleton graph is determined based on the second coordinate observation value of the joint connection vector and the parent joint associated with the joint connection vector in the next three-dimensional skeleton graph adjacent to the three-dimensional skeleton graph in which it is located.

[0084] The skeleton of human posture recognition includes several joints. The joint hierarchy is distributed from the heel joint in the center of the human body to the limbs. Each connection links the parent joint to the child joint to form a skeleton. The observation value of the spatial position of each joint can be read through the reading function, or the predicted value of the spatial position of each joint can be calculated through the spatial position and direction of the parent joint of each joint, and the connection vector from the parent joint position to the current joint position. For example, the observation value of the position of the left knee joint (KNEE_LEFT) can be directly read, and it can be calculated through the position and direction of its parent joint, the left hip joint (HIP_LEFT), and the vector from the left hip joint (HIP_LEFT) to the knee joint (KNEE_LEFT). Suppose at time t, the position of the current joint is The position of the parent joint of the current joint is The parent joint orientation expressed as a normalized quaternion is The connection vector from the parent joint position to the current joint position is Where t-1 represents the moment before time t.

[0085] In an alternative embodiment, the second coordinate observation value includes the second direction coordinate observation value. According to the joint connection vector and the second coordinate observation value of the parent joint associated with the joint connection vector in the next three-dimensional skeleton map adjacent to the three-dimensional skeleton map where the parent joint is located, the predicted value of the second position coordinate of the corresponding joint in the next adjacent three-dimensional skeleton map is determined. It may be to calculate the rotation matrix corresponding to the second direction coordinate observation value of the parent joint of the joint in the next three-dimensional skeleton map adjacent to the three-dimensional skeleton map where the parent joint is located; determine the product of the rotation matrix and the joint connection vector; add the product to the second position coordinate observation value of the parent joint of the joint in the next three-dimensional skeleton map adjacent to the three-dimensional skeleton map where the parent joint is located to obtain the predicted value of the second position coordinate of the joint in the next adjacent three-dimensional skeleton map.

[0086] The second direction coordinate observation value represented by quaternion of the parent joint in the next three-dimensional skeleton map adjacent to the three-dimensional skeleton map where the parent joint is located. Exemplarily, the quaternion q is converted into a rotation matrix R through a formula such as the Rodrigues formula, that is, the direction of the parent joint represented by quaternion The corresponding rotation matrix is

[0087] Define the state transition model And calculate that the predicted value of the second coordinate of the current joint is Then there is:

[0088]

[0089] S270. Perform coordinate fusion based on the second position coordinate observation value and the predicted value of the second position coordinate of each joint in each three-dimensional skeleton map, and update the corresponding three-dimensional skeleton map based on the coordinate fusion result.

[0090] According to the second position coordinate observation value and the predicted value of the second position coordinate of each joint in each three-dimensional skeleton map, perform coordinate fusion through the Kalman filter algorithm, and update the corresponding three-dimensional skeleton map based on the coordinate fusion result.

[0091] Exemplarily, establish the system observation equation and the state transition equation of the Kalman filter algorithm. Let the human joint position coordinate reading function be H(t), the observation noise be v(t), and the joint space position observation vector corresponding to the second position coordinate observation value be x(t); the observation equation is defined as:

[0092] x(t) = H(t) + v(t).

[0093] State transition model [[ID=--32]] Let the process noise be w(t), and the joint space position state vector corresponding to the predicted value of the second position coordinate be y(t); the state transition equation is defined as:

[0094] y(t) = F(x(t - 1)) + w(t).

[0095] According to the observation equation and state transition equation of the above system, the Kalman filtering algorithm is used for iterative optimization of prediction and update, and the optimal estimation of the system state is obtained to get the coordinate fusion result. The Kalman filtering algorithm can be the unscented Kalman filtering algorithm, the extended Kalman filtering algorithm, etc. The type of the Kalman filtering algorithm is not limited in this embodiment.

[0096] The technical solution of this embodiment is to obtain the infrared image data and depth image data of the target object, where the infrared image data and depth image data are a pair of sequence image data synchronously collected within a preset time period; input the infrared image data into a pre-trained object structure recognition model to obtain the object structure recognition result, and obtain the object structure segmentation map based on the object structure recognition result; perform data fitting on the object structure segmentation map and the corresponding depth image data to obtain the three-dimensional skeleton sequence map of the target object within the preset time period; obtain the first coordinate observation value of each joint in each three-dimensional skeleton map in the camera coordinate system in the three-dimensional skeleton sequence map; perform conversion on the first coordinate observation value to obtain the second coordinate observation value of each joint in the object coordinate system, where the second coordinate observation value includes the second position coordinate observation value; perform coordinate prediction based on the second coordinate observation value to obtain the second position coordinate prediction value of the corresponding joint in the next adjacent three-dimensional skeleton map; perform coordinate fusion according to the second position coordinate observation value and the second position coordinate prediction value of each joint in each three-dimensional skeleton map, and update the corresponding three-dimensional skeleton map based on the coordinate fusion result. The technical solution of the embodiment of the present invention solves the problems of inaccurate current pose recognition and limitations in recognition direction and angle. By combining two-dimensional and three-dimensional image data and a neural network model, a three-dimensional skeleton sequence map can be generated, which can recognize poses in any direction and angle, improve the accuracy and stability of pose recognition, and by converting the original camera coordinate system to the object coordinate system, establish the system observation equation and state transition equation for human pose recognition, and update the three-dimensional skeleton map by fusing the coordinate observation value and the prediction value, so as to improve the accuracy, stability and real-time performance of the pose recognition technology, and further improve the usability of the human pose recognition technology in assisted rehabilitation training.

[0097] Figure 4 It is a flowchart of a pose recognition method provided by an embodiment of the present invention. This embodiment and the pose recognition method in the above embodiment belong to the same inventive concept, and further describes the process of determining the target object. This method can be executed by a pose recognition device, and the device can be implemented in a software and / or hardware manner and integrated in a computer device with application development functions.

[0098] As Figure 4As shown in the figure, the gesture recognition method of this embodiment includes the following steps:

[0099] S310. Recognize the limb movements of at least one object within the image data acquisition field of view.

[0100] During the use of the system, there may be multiple objects in the camera's field of view. In the multi-person state, the system cannot recognize a specific single user. Therefore, it is necessary for the user to perform specific limb movements such as raising the hand or lifting the leg to lock a specific user among multiple people, that is, the system obtains the body identification number (body_ID) of the specific user in use.

[0101] S320. Determine the target object among at least one object according to the gesture, amplitude, and duration of the limb movement.

[0102] As Figure 5 shown, an example is given in the way that the gesture is the user raising the hand, the amplitude is the hand raising exceeding the preset angle, and the duration is 2 seconds. The locking method is as follows: after entering the hand-raising recognition, the system first initializes the body identification record, that is, the body_ID record and the current time, and then enters a loop. First, it judges whether there is a body_ID record. If there is a record, it calculates and obtains the human skeleton corresponding to the body_ID. If not, it obtains the human skeleton and the body_ID, and then calculates the raising angle of the arm. If the raising angle of the arm is greater than the preset angle, it further judges whether the body_ID and the time have been recorded. If the body_ID has been recorded, it calculates the duration of raising the hand. If the body_ID has not been recorded, it records the current body_ID and the time. On the contrary, if the raising angle of the arm does not exceed the preset angle, such as 100 degrees, it clears the current body_ID record and the time. The system displays the current image to facilitate the user to adjust the gesture. Finally, it judges whether the hand-raising time exceeds 2 seconds. If it exceeds 2 seconds, it tells the client that the body has been locked and returns the locked body body_ID, and exits the loop. If the hand-raising time does not exceed two seconds, it means that the hand-raising time is not enough or the raising angle of the arm is not enough, and the hand-raising recognition continues.

[0103] S330. Obtain the infrared image data and depth image data of the target object.

[0104] Among them, the infrared image data and the depth image data are a pair of sequence image data synchronously acquired within a preset duration.

[0105] S340. Input the infrared image data into the pre-trained object structure recognition model to obtain the object structure recognition result, and obtain the object structure segmentation map based on the object structure recognition result.

[0106] S350. Fit the object structure segmentation map with the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object within a preset time period.

[0107] S360. Obtain the first coordinate observation values of each joint in each three-dimensional skeleton map in the camera coordinate system in the three-dimensional skeleton sequence map.

[0108] S370. Convert the first coordinate observation values to obtain the second coordinate observation values of each joint in the object coordinate system.

[0109] Among them, the second coordinate observation values include the second position coordinate observation values.

[0110] S380. Perform coordinate prediction based on the second coordinate observation values to obtain the predicted second position coordinates of the corresponding joint in the next adjacent three-dimensional skeleton map.

[0111] S390. Perform coordinate fusion according to the second position coordinate observation values and the predicted second position coordinates of each joint in each three-dimensional skeleton map, and update the corresponding three-dimensional skeleton map based on the coordinate fusion result.

[0112] The technical solution of this embodiment is to identify the limb movements of at least one object within the image data acquisition field of view; determine the target object among at least one object according to the action posture, action amplitude, and action duration of the limb movement; obtain the infrared image data and depth image data of the target object, where the infrared image data and depth image data are a pair of sequence image data synchronously acquired within a preset duration; input the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtain an object structure segmentation map based on the object structure recognition result; perform data fitting on the object structure segmentation map and the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object within the preset duration; obtain the first coordinate observation value of each joint in each three-dimensional skeleton map in the camera coordinate system in the three-dimensional skeleton sequence map; perform conversion on the first coordinate observation value to obtain the second coordinate observation value of each joint in the object coordinate system, where the second coordinate observation value includes a second position coordinate observation value; perform coordinate prediction based on the second coordinate observation value to obtain the second position coordinate prediction value of the corresponding joint in the adjacent next three-dimensional skeleton map; perform coordinate fusion according to the second position coordinate observation value and the second position coordinate prediction value of each joint in each three-dimensional skeleton map, and update the corresponding three-dimensional skeleton map based on the coordinate fusion result. The technical solution of the embodiment of the present invention solves the problems of inaccurate current posture recognition and limitations in recognition direction and angle. It can generate a three-dimensional skeleton sequence map by combining two-dimensional and three-dimensional image data and a neural network model, can recognize postures in any direction and angle, improve the accuracy and stability of posture recognition, and determine the target object from a multi-person scenario by recognizing limb movements, thereby improving the usability of posture recognition in rehabilitation training.

[0113] Figure 6 FIG. is a schematic structural diagram of a posture recognition device provided by an embodiment of the present invention. This embodiment is applicable to the scenario of posture recognition. The posture recognition device can be implemented in a software and / or hardware manner and integrated into a computer terminal device with application development functions.

[0114] As Figure 6 shown, the posture recognition device includes: an image acquisition module 410, an object structure segmentation map generation module 420, and a three-dimensional skeleton sequence map generation module 430.

[0115] Among them, the image acquisition module 410 is used to acquire the infrared image data and depth image data of the target object, where the infrared image data and depth image data are a pair of sequence image data obtained by synchronous acquisition within a preset time period; the object structure segmentation map generation module 420 is used to input the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtain an object structure segmentation map based on the object structure recognition result; the three-dimensional skeleton sequence map generation module 430 is used to perform data fitting on the object structure segmentation map and the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object within a preset time period.

[0116] The technical solution of this embodiment, by acquiring the infrared image data and depth image data of the target object, where the infrared image data and depth image data are a pair of sequence image data obtained by synchronous acquisition within a preset time period; inputting the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtaining an object structure segmentation map based on the object structure recognition result; performing data fitting on the object structure segmentation map and the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object within a preset time period. The technical solution of the embodiment of the present invention solves the problems of inaccurate current pose recognition and limitations in recognition direction and angle. It can generate a three-dimensional skeleton sequence map by combining two-dimensional and three-dimensional image data and a neural network model, and can recognize poses in any direction and angle, improving the accuracy and stability of pose recognition.

[0117] In an alternative embodiment, the device further includes:

[0118] The three-dimensional skeleton map update module is used to obtain the first coordinate observation value of each joint in each three-dimensional skeleton map in the camera coordinate system in the three-dimensional skeleton sequence map; perform a conversion on the first coordinate observation value to obtain the second coordinate observation value of each joint in the object coordinate system, where the second coordinate observation value includes a second position coordinate observation value; perform coordinate prediction based on the second coordinate observation value to obtain the second position coordinate prediction value of the corresponding joint in the next adjacent three-dimensional skeleton map; perform coordinate fusion according to the second position coordinate observation value and the second position coordinate prediction value of each joint in each three-dimensional skeleton map, and update the corresponding three-dimensional skeleton map based on the coordinate fusion result.

[0119] In an alternative embodiment, the three-dimensional skeleton map update module is specifically used for:

[0120] Determine the relative position relationship between the coordinate reference joint in each three-dimensional skeleton diagram and the origin in the object coordinate system; according to the first coordinate observation value of the coordinate reference joint in the camera coordinate system and the relative position relationship, determine the coordinate system mapping relationship between the camera coordinate system and the object coordinate system; according to the coordinate system mapping relationship and the position relationship between the non-coordinate reference joint and the corresponding coordinate reference joint in each three-dimensional skeleton diagram, convert the first coordinate observation value of the non-coordinate reference joint to obtain the second coordinate observation value of each joint in the object coordinate system for each three-dimensional skeleton diagram.

[0121] In an alternative embodiment, the three-dimensional skeleton diagram update module is further configured to:

[0122] Determine the joint connection vector between each joint and the corresponding parent joint of the joint according to the second coordinate observation values of each joint and the corresponding parent joint; determine the predicted second position coordinate value of the corresponding joint in the next adjacent three-dimensional skeleton diagram according to the joint connection vector and the second coordinate observation value of the parent joint associated with the joint connection vector in the next adjacent three-dimensional skeleton diagram of its own three-dimensional skeleton diagram.

[0123] In an alternative embodiment, the second coordinate observation value includes the second direction coordinate observation value, and the three-dimensional skeleton diagram update module is further configured to:

[0124] Calculate the rotation matrix corresponding to the second direction coordinate observation value of the parent joint of the joint in the next adjacent three-dimensional skeleton diagram of its own three-dimensional skeleton diagram; determine the product of the rotation matrix and the joint connection vector; add the product to the second position coordinate observation value of the parent joint of the joint in the next adjacent three-dimensional skeleton diagram of its own three-dimensional skeleton diagram to obtain the predicted second position coordinate value of the joint in the next adjacent three-dimensional skeleton diagram.

[0125] In an alternative embodiment, the apparatus further includes:

[0126] A target object determination module, configured to identify the limb movements of at least one object within the image data acquisition field of view; determine the target object among at least one object according to the action posture, action amplitude, and action duration of the limb movement.

[0127] In an alternative embodiment, the apparatus further includes:

[0128] An outlier rejection module, configured to compare and analyze the first coordinate observation value with a pre-determined sample coordinate observation value to obtain a comparison and analysis result; determine the outliers in the first coordinate observation value according to the comparison and analysis result, and reject the outliers; wherein, the sample coordinate observation value is obtained by synchronously collecting infrared image data and depth image data of a sample object in a static and moving state and performing coordinate reading within a preset sample coordinate acquisition duration.

[0129] The pose recognition device provided by an embodiment of the present invention can execute the pose recognition method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0130] Figure 7 It is a schematic structural diagram of a pose recognition system provided by an embodiment of the present invention.

[0131] As Figure 7 shown, the pose recognition system includes:

[0132] A three-dimensional depth camera 510, an infrared camera 520, and a computing and processing platform 530;

[0133] Among them, the three-dimensional depth camera 510 is used to collect depth image data of a target object within a preset time;

[0134] The infrared camera 520 is used to collect infrared image data of the target object within a preset time;

[0135] The computing and processing platform 530 is used to perform pose recognition based on the depth image data and the infrared image data according to the method provided by any embodiment of the present invention.

[0136] The technical solution of this embodiment uses a three-dimensional depth camera to collect depth image data of a target object within a preset time; an infrared camera to collect infrared image data of the target object within a preset time; and a computing and processing platform to perform pose recognition based on the depth image data and the infrared image data according to the method provided by any embodiment of the present invention. The technical solution of the embodiment of the present invention solves the problems of inaccurate current pose recognition and limitations in recognition direction and angle. By combining two-dimensional and three-dimensional image data and a neural network model, a three-dimensional skeleton sequence diagram can be generated, which can recognize poses in any direction and angle, improving the accuracy and stability of pose recognition.

[0137] To ensure the real-time performance and usability of the pose recognition system, this embodiment deploys the pose recognition method to the computing and processing platform, which can be an edge computing platform. This embodiment preferably runs the inference of the pre-trained object structure recognition model on a GPU (Graphics Processing Unit) for parallel computing inference, thereby improving the system's computing speed and real-time performance. Taking the configuration of the edge computing platform shown in Table 2 as an example, the system's computing speed expressed in frames per second (FPS) is not less than 25 FPS, that is, the system's computing bandwidth is 25 Hz.

[0138] Table 2

[0139] Configuration Item Description Operating System Windows 11 64-bit 16G System Memory 512G

[0140] In a preferred embodiment of the present invention, as Hard Disk shown, a three-dimensional depth camera composed of a three-dimensional depth sensor based on the principle of time-of-flight (ToF), a color camera, and an infrared camera is used as the sensor, and an edge computing platform with a GPU (Graphics Processing Unit) is used as the computing and processing platform; first, the infrared image data and depth image data of the camera are read through the software development interface (API, Application Programming Interface) of the three-dimensional depth camera, and the two-dimensional human joints and the segmentation map of each human part are obtained respectively using the infrared image data through a pre-trained human part recognition model. At the same time, the depth image data is fitted with the two-dimensional human joints to obtain the human joints with three-dimensional coordinates and form a skeleton.

[0141] Figure 8 It is a schematic structural diagram of a rehabilitation mirror provided by an embodiment of the present invention.

[0142] As Figure 9 shown, the rehabilitation mirror includes:

[0143] an image acquisition module 610, a calculation and processing module 620, and a display screen 630.

[0144] Among them, the image acquisition module 610 is used to acquire the depth image data and infrared image data of the target object within a preset time; the calculation and processing module 620 is used to perform pose recognition and generate a three-dimensional skeleton sequence diagram and / or a three-dimensional skeleton diagram based on the depth image data and the infrared image data according to the pose recognition method provided by any embodiment of the present invention; the display screen 630 is used to display the three-dimensional skeleton sequence diagram and / or the three-dimensional skeleton diagram.

[0145] The patient can observe the three-dimensional skeleton sequence diagram and / or the three-dimensional skeleton diagram on the display screen 630 to understand whether his actions are correct and the areas that need to be improved, so as to adjust the actions in time and improve the effect of rehabilitation training.

[0146] In the technical solution of this embodiment, an image acquisition module is used to acquire depth image data and infrared image data of a target object within a preset time; a calculation and processing module is used to perform pose recognition based on the depth image data and the infrared image data according to the pose recognition method provided in any embodiment of the present invention, and generate a three-dimensional skeleton sequence diagram and / or a three-dimensional skeleton diagram; a display screen is used to display the three-dimensional skeleton sequence diagram and / or the three-dimensional skeleton diagram. The technical solution of the embodiment of the present invention solves the problem that the pose recognition applied to the rehabilitation mirror is inaccurate at present, and the recognition direction and angle are limited, resulting in more difficulties for patients with pose problems when using the rehabilitation mirror for training and a poor user experience. By combining two-dimensional and three-dimensional image data and a neural network model, a three-dimensional skeleton sequence diagram can be generated, which can recognize poses in any direction and angle, reduce the motion requirements during the training of patients, and improve the user experience of patient users.

[0147] After sports injury repair, effective rehabilitation training is required to ensure complete injury recovery. With the development of computer technology, rehabilitation mirrors are increasingly used in medical rehabilitation training. In particular, computer vision technology has played an important role in the monitoring of rehabilitation training, and the computer vision technology for rehabilitation training monitoring first needs to solve the problem of human pose recognition.

[0148] In this embodiment, the rehabilitation mirror can be a kind of medical rehabilitation assistance device with pose recognition technology, which can provide efficient and accurate support for the rehabilitation training of patients.

[0149] Optionally, the above-mentioned rehabilitation mirror further includes:

[0150] A rehabilitation plan customization module, which is used to determine a target rehabilitation training plan according to the rehabilitation requirement data of the target object.

[0151] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present invention. Figure 10 The block diagram of an exemplary computer device 12 suitable for implementing the embodiments of the present invention is shown. Figure 10 The displayed computer device 12 is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities, such as intelligent controllers, servers, mobile phones and other terminal devices.

[0152] As Figure 10 shown, the computer device 12 is presented in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0153] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor, or a local bus using any of a variety of bus architectures. By way of example, such architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0154] Computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by computer device 12, including volatile and nonvolatile media, removable and non-removable media.

[0155] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer device 12 can further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing on non-removable, nonvolatile magnetic media ( Figure 10 not shown, typically called a "hard disk drive"). Although Figure 10 not shown in the figures, a disk drive for reading and writing on removable nonvolatile disks (such as a "floppy disk"), and an optical disk drive for reading and writing on removable nonvolatile optical disks (such as a CD-ROM, DVD-ROM, or other optical media) can be provided. In these cases, each drive can be connected to bus 18 by one or more data media interfaces. System memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the present invention.

[0156] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in system memory 28, and such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a network environment. Program modules 42 generally carry out the functions and / or methods described in the embodiments of the present invention.

[0157] The computer device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the computer device 12, and / or communicate with any device that enables the computer device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Moreover, the computer device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the computer device 12 through the bus 18. It should be understood that although Figure 10 Figure 10 not shown in the figure, other hardware and / or software modules can be used in combination with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0158] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, for example, implementing the posture recognition method provided by the embodiments of the present invention. The method includes:

[0159] Obtaining infrared image data and depth image data of a target object, where the infrared image data and the depth image data are a pair of sequence image data synchronously collected within a preset time period;

[0160] Inputting the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtaining an object structure segmentation map based on the object structure recognition result;

[0161] Performing data fitting on the object structure segmentation map and the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object within a preset time period.

[0162] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the posture recognition method provided by any embodiment of the present invention. The method includes:

[0163] Obtaining infrared image data and depth image data of a target object, where the infrared image data and the depth image data are a pair of sequence image data synchronously collected within a preset time period;

[0164] Inputting the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtaining an object structure segmentation map based on the object structure recognition result;

[0165] Perform data fitting on the object structure segmentation diagram and the corresponding depth image data to obtain the three-dimensional skeleton sequence diagram of the target object within a preset time period.

[0166] The computer storage medium of the embodiments of the present invention may adopt any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.

[0167] The computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium may send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.

[0168] The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0169] The computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, Python, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0170] An embodiment of the present invention also provides a computer program product, including a computer program which, when executed by a processor, implements the gesture recognition method provided in any embodiment of the present application.

[0171] In the process of implementing the computer program product, the computer program code for performing the operations of the present invention can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, Python, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or can be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0172] Those of ordinary skill in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they can be implemented with program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0173] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A gesture recognition method, characterized in that, Including: Obtain the infrared image data and depth image data of the target object, where the infrared image data and the depth image data are a pair of sequence image data synchronously collected within a preset time period; Input the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and obtain an object structure segmentation map based on the object structure recognition result; Perform data fitting on the object structure segmentation map and the corresponding depth image data to obtain a three-dimensional skeleton sequence map of the target object within the preset time period; Obtain the first coordinate observation value of each joint in each three-dimensional skeleton map in the three-dimensional skeleton sequence map in the camera coordinate system; Convert the first coordinate observation value to obtain the second coordinate observation value of each joint in the object coordinate system, where the second coordinate observation value includes the second position coordinate observation value; Perform coordinate prediction based on the second coordinate observation value to obtain the second position coordinate prediction value of the corresponding joint in the next adjacent three-dimensional skeleton map; Perform coordinate fusion according to the second position coordinate observation value and the second position coordinate prediction value of each joint in each three-dimensional skeleton map, and update the corresponding three-dimensional skeleton map based on the coordinate fusion result.

2. The method according to claim 1, wherein The conversion of the first coordinate observation value to obtain the second coordinate observation value of each joint in the object coordinate system includes: Determine the relative position relationship between the coordinate reference joint in each three-dimensional skeleton map and the origin in the object coordinate system; Determine the coordinate system mapping relationship between the camera coordinate system and the object coordinate system according to the first coordinate observation value of the coordinate reference joint in the camera coordinate system and the relative position relationship; According to the coordinate system mapping relationship and the position relationship between the non-coordinate reference joint and the corresponding coordinate reference joint in each three-dimensional skeleton map, convert the first coordinate observation value of the non-coordinate reference joint to obtain the second coordinate observation value of each joint in each three-dimensional skeleton map in the object coordinate system.

3. The method according to claim 1, characterized in that, The coordinate prediction based on the second coordinate observation value to obtain the second position coordinate prediction value of the corresponding joint in the next adjacent three-dimensional skeleton map includes: Determine the joint connection vector between each joint and the corresponding parent joint according to the second coordinate observation value of each joint and the corresponding parent joint; Determine the second position coordinate prediction value of the corresponding joint in the next adjacent three-dimensional skeleton map according to the joint connection vector and the second coordinate observation value of the parent joint associated with the joint connection vector in the next adjacent three-dimensional skeleton map of its own three-dimensional skeleton map.

4. The method according to claim 3, characterized in that The second coordinate observation value includes the second direction coordinate observation value. The determination of the second position coordinate prediction value of the corresponding joint in the next adjacent three-dimensional skeleton map according to the joint connection vector and the second coordinate observation value of the parent joint associated with the joint connection vector in the next adjacent three-dimensional skeleton map of its own three-dimensional skeleton map includes: Calculate the rotation matrix corresponding to the second direction coordinate observation value of the parent joint of the joint in the next adjacent three-dimensional skeleton map of its own three-dimensional skeleton map; determining the product of the rotation matrix and the joint connection vector; The product is added to the second position coordinate observation value of the parent joint of the joint in the next three-dimensional skeleton graph adjacent to the three-dimensional skeleton graph where the joint is located, to obtain the second position coordinate prediction value of the joint in the next adjacent three-dimensional skeleton graph.

5. The method according to claim 1, wherein Before acquiring infrared image data and depth image data of the target object, the following steps are also included: identifying a body movement of at least one subject within a field of view of image data acquisition; A target object is determined in the at least one object according to the movement posture, movement amplitude and movement duration of the limb movement.

6. The method according to claim 1, wherein Before converting the first coordinate observation value, the method further includes: Comparing and analyzing the first coordinate observation value with a predetermined sample coordinate observation value to obtain a comparative analysis result; Determine outliers in the first coordinate observation values according to the comparative analysis results, and remove the outliers; The sample coordinate observation values are obtained by synchronously collecting infrared image data and depth image data of the sample object in static and moving states and reading the coordinates within a preset sample coordinate collection time.

7. A posture recognition device, characterized in that, include: An image acquisition module is used to acquire infrared image data and depth image data of a target object, wherein the infrared image data and the depth image data are a pair of sequential image data acquired synchronously within a preset time period; an object structure segmentation map generation module, configured to input the infrared image data into a pre-trained object structure recognition model to obtain an object structure recognition result, and to obtain an object structure segmentation map based on the object structure recognition result; a three-dimensional skeleton sequence diagram generation module, configured to perform data fitting on the object structure segmentation diagram and the corresponding depth image data to obtain a three-dimensional skeleton sequence diagram of the target object within the preset time length; A three-dimensional skeleton graph updating module is used to obtain the first coordinate observation value of each joint in each three-dimensional skeleton graph in the three-dimensional skeleton sequence graph in the camera coordinate system; Converting the first coordinate observation value to obtain a second coordinate observation value of each joint in an object coordinate system, wherein the second coordinate observation value includes a second position coordinate observation value; Performing coordinate prediction based on the second coordinate observation value to obtain a second position coordinate prediction value of the corresponding joint in the next adjacent three-dimensional skeleton image; Coordinate fusion is performed according to the second position coordinate observation value and the second position coordinate prediction value of each joint in each three-dimensional skeleton image, and the corresponding three-dimensional skeleton image is updated based on the coordinate fusion result.

8. A posture recognition system, characterized in that, include: 3D depth camera, infrared camera and computing processing platform; Wherein, the three-dimensional depth camera is used to collect depth image data of the target object within a preset time; The infrared camera is used to collect infrared image data of the target object within a preset time; The computing and processing platform is used to perform gesture recognition based on the depth image data and the infrared image data according to the gesture recognition method according to any one of claims 1 to 6.

9. A computer device, characterized in that, The computer device comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the gesture recognition method according to any one of claims 1-6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the gesture recognition method according to any one of claims 1-6.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the gesture recognition method according to any one of claims 1-6.

12. A rehabilitation mirror, characterized in that, Comprising: An image acquisition module, a calculation and processing module, and a display screen; Wherein, the image acquisition module is configured to acquire depth image data and infrared image data of a target object within a preset time; The calculation and processing module is configured to perform gesture recognition and generate a three-dimensional skeleton sequence diagram and / or a three-dimensional skeleton diagram based on the depth image data and the infrared image data according to the method according to any one of claims 1-6; The display screen is configured to display the three-dimensional skeleton sequence diagram and / or the three-dimensional skeleton diagram.

Citation Information

Patent Citations

  • Intelligent desk with sitting posture correcting function and correcting method implemented by intelligent desk

    CN103908066A

  • Behavior and identity joint recognition method based on human skeleton sequence

    CN108764107A

  • Skeleton detection tracking method and device and electronic equipment

    CN114419103A

  • Augmented reality-based navigation processing method, apparatus, device and program product

    CN116242370A