Robot-based gait monitoring method, device, equipment and storage medium
By implementing a gait monitoring method on the robot, using the identity feature extraction model to extract and fuse the identity features, the problem of the existing robots requiring close distance recognition is solved, the convenience and accuracy of recognition is improved, and the degree of intelligence of the robot is improved.
Patent Information
- Application Number
- CN202210951063.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-09
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-08-09
AI Technical Summary
Existing robots need close-range recognition when identifying people, resulting in a low degree of intelligence.
By implementing a gait monitoring method on the robot, the human silhouette sequence and key point sequence in the captured video are obtained, and the silhouette feature extraction network, key point feature extraction network and multimodal feature hybrid network in the identity feature extraction model are used to extract and fuse identity features for identification.
It improves the convenience and accuracy of the robot's personnel identity recognition, reduces the requirements for identification distance, and improves the intelligence of the robot.
Smart Images

Figure CN115439927B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular, to a gait monitoring method, device, equipment, and storage medium based on a robot. Background Art
[0002] With the development of artificial intelligence technology, robots can provide various services such as cleaning, entertainment, and communication, bringing convenience to people's lives and work.
[0003] In related technologies, a robot uses face recognition technology to perform image recognition on a user's face image to determine the user's identity, or the robot uses pupil recognition technology to recognize the user's pupils to determine the user's identity. Subsequently, services are provided to users whose identities meet the requirements.
[0004] In the above methods, the robot needs to perform close-range recognition on the user's face, and the user needs to approach the robot to be recognized. Therefore, the intelligence level of the robot still needs to be improved. Summary of the Invention
[0005] Embodiments of the present disclosure provide a gait monitoring method, device, equipment, and storage medium based on a robot to improve the intelligence level of the robot.
[0006] In a first aspect, embodiments of the present disclosure provide a gait monitoring method based on a robot, including:
[0007] Obtaining a captured video of the robot;
[0008] Performing image processing on video frames of the captured video to obtain a human silhouette sequence and a human key point sequence of a person in the captured video;
[0009] Extracting features from the human silhouette sequence through a silhouette feature extraction network in an identity feature extraction model to obtain a human silhouette feature of the person;
[0010] Extracting features from the human key point sequence through a key point feature extraction network in the identity feature extraction model to obtain a human key point feature of the person;
[0011] Performing feature fusion on the human silhouette feature and the human key point feature through a multi-modal feature mixing network in the identity feature extraction model to obtain an identity feature of the person;
[0012] Performing identity recognition on the person according to the identity feature.
[0013] In a second aspect, embodiments of the present disclosure provide a gait monitoring device based on a robot, including:
[0014] A video acquisition unit, configured to acquire a captured video of the robot;
[0015] A video processing unit, configured to perform image processing on video frames of the captured video to obtain a sequence of human silhouettes and a sequence of human key points of the people in the captured video;
[0016] A silhouette feature extraction unit, configured to extract features from the sequence of human silhouettes through a silhouette feature extraction network in an identity feature extraction model to obtain human silhouette features of the people;
[0017] A key point feature extraction unit, configured to extract features from the sequence of human key points through a key point feature extraction network in the identity feature extraction model to obtain human key point features of the people;
[0018] A feature fusion unit, configured to fuse the human silhouette features and the human key point features through a multi-modal feature mixing network in the identity feature extraction model to obtain identity features of the people;
[0019] An identity recognition unit, configured to perform identity recognition on the people according to the identity features.
[0020] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory; the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the gait monitoring method based on a robot as described in the first aspect above.
[0021] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored, and when a processor executes the computer execution instructions, the gait monitoring method based on a robot as described in the first aspect above is implemented.
[0022] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, which includes computer execution instructions, and when a processor executes the computer execution instructions, the gait monitoring method based on a robot as described in the first aspect above is implemented.
[0023] The gait monitoring method, device, equipment and storage medium based on a robot provided by the embodiments of the present disclosure propose a method for personnel identification through gait monitoring on the robot. Compared with other identification methods, this method has lower requirements for the distance between the personnel and the robot and is more convenient. During the process of personnel identification, the embodiments of the present disclosure utilize an identity feature extraction network including a silhouette feature extraction network, a key point feature extraction network, and a multi-modal feature mixing network to identify the identity features of personnel from the captured video. The identity features fuse the human silhouette and human key points that can reflect the gait of the personnel, improving the accuracy of the identity features, the accuracy of personnel identification using the gait monitoring method on the robot, and the intelligence level of the robot. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0025] Figure 1 Schematic flowchart of the gait monitoring method based on a robot provided by the embodiments of the present disclosure Figure 1 ;
[0026] Figure 2 Schematic structure of the identity feature extraction model provided by the embodiments of the present disclosure Figure 1 ;
[0027] Figure 3 Schematic structure of the body feature extraction model provided by the embodiments of the present application Figure 2 ;
[0028] Figure 4 Schematic flowchart of the gait monitoring method based on a robot provided by the embodiments of the present disclosure Figure 2 ;
[0029] Figure 5 Schematic diagram of the structure of the two-dimensional residual network provided by the embodiments of the present application;
[0030] Figure 6 Schematic flowchart of the gait monitoring method based on a robot provided by the embodiments of the present disclosure Figure 3 ;
[0031] Figure 7 Schematic diagram of the structure of the health monitoring network provided by the embodiments of the present application;
[0032] Figure 8Schematic diagram of the data processing flow of the leg type health monitoring network provided by the embodiment of the present application;
[0033] Figure 9 Block diagram of the structure of the gait monitoring device based on a robot provided by the embodiment of the present disclosure;
[0034] Figure 10 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present disclosure. Detailed implementation manners
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.
[0036] In the related art, when a robot identifies a person's identity, it usually needs to identify at close range, such as identifying a person's facial features, voice, or iris and other human biological features at close range.
[0037] Considering that a person's gait is also recognizable and the requirement for the recognition distance of gait recognition is relatively low, the embodiments of the present disclosure provide a gait monitoring method, device, equipment, and storage medium based on a robot. In the embodiments of the present disclosure, a captured video of the robot is obtained, a sequence of human silhouettes and a sequence of human key points of a person are obtained from the captured video, and based on the sequence of human silhouettes, the sequence of human key points, and an identity feature extraction model, the identity of the person is recognized by combining the silhouette features and key point features of the person. In this way, gait recognition is applied to the robot, improving the convenience and accuracy of the robot's person identity recognition, enhancing the intelligence level of the robot, and thus improving the user experience.
[0038] Among them, the execution subject of the embodiments of the present disclosure may be a robot; alternatively, the execution subject of the embodiments of the present disclosure may be a server connected to the robot. The robot sends the captured video to the server, and the server performs identity recognition of the person based on the captured video and then returns the recognition result to the robot.
[0039] Figure 1 Schematic flow of the gait monitoring method based on a robot provided by the embodiment of the present disclosure Figure 1 As Figure 1 shown, the gait monitoring method based on a robot includes:
[0040] S101. Obtain a captured video of the robot.
[0041] In this embodiment, a camera is provided on the robot. The robot can perform video shooting through the camera with the user's authorization to obtain a captured video.
[0042] In one example, the robot performs real-time video shooting to obtain a captured video.
[0043] In another example, the robot can perform video shooting to obtain a captured video when receiving a service request from the user (such as a login request or a request for providing entertainment services).
[0044] In another example, the robot can perform video shooting when detecting the appearance of a person to obtain a captured video to ensure that there is a person in the captured video.
[0045] S102. Perform image processing on the video frames of the captured video to obtain a sequence of human silhouettes and a sequence of human key points of the people in the captured video.
[0046] Among them, a silhouette can be understood as a contour image. The sequence of human silhouettes includes multiple human silhouettes, and a human silhouette can be understood as a human contour image.
[0047] Among them, the sequence of human key points includes multiple groups of human key points, and each group of human key points includes the image positions of multiple human key points of the people in the video frame. Based on one video frame, a human silhouette and a group of human key points of the people appearing in the video frame can be obtained. Therefore, based on the captured video, a sequence of human silhouettes and a sequence of human key points of the people in the captured video can be obtained.
[0048] The human silhouette of a person can roughly reflect the body contour of the person in the video frame, and the body contour is closely related to the body posture. Therefore, the sequence of human silhouettes of a person can reflect the dynamic changes of the body posture of the people in the captured video, especially the gait of the people. The human key points of a person can reflect the body posture of the human body in the video frame in detail. Therefore, the sequence of human key points of a person can reflect the dynamic changes of the body posture of the people in the captured video, especially the gait of the people.
[0049] In this embodiment, taking pedestrians as the target, image processing such as target tracking, target segmentation, and key point detection can be performed on the video frames of the captured video to obtain the human silhouette and human key points of the people in the video frame. In this way, the human silhouettes and human key points of the people in multiple video frames are obtained. According to the order of the video frames in the captured video, the human silhouettes of the people in multiple video frames can be combined to obtain a sequence of human silhouettes, and the human key points of the people in multiple video frames can be combined to obtain a sequence of human key points.
[0050] S103. Use the silhouette feature extraction network in the identity feature extraction model to extract features from the human silhouette sequence, and obtain the human silhouette features of the person.
[0051] Among them, the identity feature extraction model is used to extract the identity features of the person in the captured video. The input data of the identity feature extraction model is the human silhouette sequence and the human key point sequence of the person, and the output data is the identity feature of the person. The identity feature extraction model can be pre-trained to obtain a trained identity feature extraction model, and then the trained identity feature extraction model can be deployed on the robot or the server communicatively connected to the robot to provide identity feature extraction services for the robot.
[0052] Such as Figure 2 ( Figure 2 Schematic diagram of the structure of the identity feature extraction model provided by the embodiment of the present disclosure Figure 1 ) As shown, the identity feature extraction model may include a silhouette feature extraction network, a key point feature extraction network, and a multi-modal feature mixing network. Among them, the silhouette feature extraction network is used to extract the silhouette features of the person based on the human silhouette sequence of the person; the key point extraction network is used to extract the human key point features of the person based on the human key point sequence of the person; the multi-modal feature mixing network is used to perform feature fusion on the human silhouette features and the human key point features of the person to obtain the identity feature of the person.
[0053] In this embodiment, after obtaining the human silhouette sequence of the person in the captured video, the human silhouette sequence of the person can be input into the silhouette feature extraction network in the identity feature extraction model. In the silhouette feature extraction network, feature extraction is performed on multiple human silhouettes in the human silhouette sequence to obtain the human silhouette features of the person. Among them, the silhouette feature extraction network can be a deep neural network.
[0054] S104. Use the key point feature extraction network in the identity feature extraction model to extract features from the human key point sequence, and obtain the human key point features of the person.
[0055] In this embodiment, after obtaining the human key point sequence of the person in the captured video, the human key point sequence can be input into the key point feature extraction network in the identity feature extraction model. In the key point feature network, feature extraction is performed on multiple groups of human key points in the human key point sequence to obtain the human key point features of the person. Among them, the key point feature extraction network can be a deep neural network.
[0056] S105. Use the multi-modal feature mixing network in the identity feature extraction model to perform feature fusion on the human silhouette features and the human key point features to obtain the identity feature of the person.
[0057] In this embodiment, the human silhouette feature and the human key point feature of a person can be input into the multi-modal feature mixing network in the identity feature extraction model. In the multi-modal feature mixing network, the human silhouette feature and the human key point feature are fused, and finally the identity feature output by the multi-modal feature mixing network is obtained. Thus, by combining two features that can reflect the gait of a person: the human silhouette feature and the human key point feature, the diversity and accuracy of the identity feature are improved.
[0058] S106. Identify the identity of the person according to the identity feature.
[0059] In this embodiment, after obtaining the identity feature of the person, the identity of the person can be identified by matching the identity feature of the person with the identity features of multiple service objects in the identity feature library, and the identity recognition result is obtained. Among them, the identity recognition result may include: the person is a service object of the robot or the person is not a service object of the robot.
[0060] In the embodiment of the present disclosure, a human silhouette sequence and a human key point sequence reflecting the characteristics of human gait are extracted from the captured video of the robot. Through an identity feature extraction model including a silhouette feature extraction network, a key point feature extraction network, and a multi-modal feature mixing network, feature extraction and feature fusion are performed on the human silhouette sequence and the human key point sequence to obtain an identity feature, which improves the diversity and accuracy of the identity feature. Identifying the identity of a person based on the identity feature improves the accuracy of person identity recognition. Thus, the convenience and accuracy of identity recognition are improved.
[0061] In some embodiments, a possible implementation manner of S102 includes: performing target tracking on the person in the captured video, determining the human image area in the video frame of the captured video; performing edge segmentation on the human image area through a target segmentation algorithm to obtain a human silhouette sequence; and performing key point detection on the human image area through a key point estimation algorithm to obtain a human key point sequence. Thus, by using target tracking to first determine the human image area, and then using the target segmentation algorithm and the key point estimation algorithm to perform human edge segmentation and key point detection in the human image area, the accuracy of human edge segmentation and key point detection is obtained, and the accuracy of the human silhouette sequence and the human key point sequence is improved.
[0062] In this implementation manner, the process of obtaining the human silhouette sequence and the human key point sequence based on the captured video is equivalent to the human body preprocessing process. In this process, a target tracking algorithm is used to track and monitor the people in the captured video, and the human body image area in the video frames of the captured video is obtained (annotated by a rectangular box in the video frame). The human body image area is input into the target segmentation algorithm, and the human body edges in the human body image area are segmented through the target segmentation algorithm to obtain the human silhouette of the person in the video frame. Based on the human silhouettes of the person in multiple video frames, the human silhouette sequence of the person is obtained. The human body image area is input into the key point estimation algorithm, and through the key point estimation algorithm, the image positions of the human key points in the human body image area are estimated to obtain the human key points of the person in the video frame. Based on the human key points of the person in multiple video frames, the human key point sequence of the person is obtained.
[0063] In one example, the target tracking algorithm can adopt the YOLOv5 tracking algorithm to improve the accuracy of person tracking in the captured video through the YOLOv5 tracking algorithm.
[0064] In one example, the key point tracking algorithm can adopt the human body two-dimensional (2D) key point estimation algorithm to accurately extract the human 2D key points of the person.
[0065] In one example, the human key points in the human key point sequence can include body key points and facial key points. Among them, the body key points can include torso key points and limb key points. For example, using the 2D key point estimation algorithm to estimate the key points, a set of key points composed of 25 torso key points, 70 facial key points, and 21 hand key points of the person is obtained. Therefore, the richness of the human key point features is improved through the human key points of multiple parts, and further the accuracy of identity recognition is improved.
[0066] In some embodiments, considering that the amount of data that the identity feature extraction model can input is limited, before performing image processing on the video frames in the captured video, the video frames in the captured video can be screened, and then image processing is performed on the screened video frames to obtain the human silhouette sequence and the human key point sequence of the person, so as to improve the quality of the video frames used for identity feature extraction and make the number of the screened video frames meet the requirements of the identity feature extraction model through the image screening method.
[0067] In one example, when screening, a continuous preset number of video frames can be screened according to the moments when the people appear in the captured video. Then, image processing is performed on the continuous preset number of video frames to obtain the human silhouette sequence and the human key point sequence of the person. Thus, it is ensured as much as possible that there are people in the screened video frames and the continuity of the screened video frames in time is ensured.
[0068] For example, the identity feature extraction model requires the number of frames of the input data to be 30, that is, the human silhouette sequence contains 30 human silhouettes of a person, and the human key point sequence contains 30 groups of human key points.
[0069] In some embodiments, such as Figure 3 ( Figure 3 is a schematic structural diagram of the body feature extraction model provided by the embodiment of the present application Figure 2 ) As shown, the silhouette feature extraction network includes a spatial domain feature extraction network and a pooling network. After the human silhouette sequence is input into the spatial domain feature extraction network, through the feature extraction of the spatial domain feature extraction network, the silhouette spatial domain feature output by the spatial feature extraction network is obtained. After the silhouette spatial domain feature is input into the pooling network, through the pooling process of the pooling network, the human silhouette feature is obtained. In this way, through the spatial domain feature extraction network and the pooling network, the accuracy of silhouette feature extraction is improved.
[0070] Among them, the spatial domain feature extraction network refers to feature extraction in the spatial dimension. Since the human silhouettes in the human silhouette sequence are two-dimensional images, which reflect more the features of the human gait in the spatial dimension, the spatial domain feature extraction network can improve the accuracy of human silhouette feature extraction.
[0071] In some embodiments, such as Figure 3 As shown, the key point feature extraction network includes a time domain feature extraction network. After the human key point sequence is input into the time domain feature extraction network, through the feature extraction of the time domain feature extraction network, the human key point feature output by the time domain feature extraction network is obtained. In this way, through the time domain feature extraction network, the accuracy of human key point feature extraction is improved.
[0072] Among them, the time domain feature extraction network refers to feature extraction in the time dimension (i.e., the time sequence dimension). Since the human key point sequence reflects the change of human key points over time, that is, the features of the human gait in the time dimension, the time domain feature extraction network can improve the accuracy of human key point feature extraction.
[0073] Figure 4 is a schematic flowchart of the gait monitoring method based on a robot provided by the embodiment of the present disclosure Figure 2 , in the embodiment of the present disclosure, the silhouette feature extraction network includes a spatial domain feature extraction network and a pooling network, and the key point feature extraction network includes a time domain feature extraction network. As Figure 4 shown, the gait monitoring method of the robot includes:
[0074] S401. Obtain the captured video of the robot.
[0075] S402. Perform image processing on the video frames of the captured video to obtain a human silhouette sequence and a human key point sequence of the person in the captured video.
[0076] Among them, the implementation principles and technical effects of S401 to S402 can be referred to the foregoing embodiments and will not be elaborated herein.
[0077] S403: Input the human silhouette sequence into the spatial domain feature extraction network, and perform feature extraction on the human silhouette sequence in the spatial domain feature extraction network to obtain the silhouette spatial domain features.
[0078] In this embodiment, as Figure 3 shown, the input data of the spatial domain feature extraction network includes the human silhouette sequence of the personnel in the captured video, and the output data includes the silhouette spatial domain features. The human silhouettes in the human silhouette sequence are two-dimensional images with spatial characteristics, and the spatial domain feature extraction network can be used to perform image feature extraction on the human silhouette sequence in the spatial dimension to obtain the silhouette spatial domain features.
[0079] In a possible implementation manner, the spatial domain feature extraction network is a two-dimensional residual network, and the convolutional layer in the two-dimensional residual network is a two-dimensional convolutional layer. Based on this, S203 includes: inputting the human silhouette sequence of the personnel into the two-dimensional residual network, and performing feature extraction on the human silhouette sequence through the two-dimensional residual network to obtain the silhouette spatial domain features. Thus, the extraction of human silhouette features is realized by using the residual network, the accuracy of human silhouette feature extraction is improved, and further the accuracy of personnel identity recognition is improved.
[0080] In one example, Figure 5 is the structural schematic diagram of the two-dimensional residual network provided by the embodiment of the present application. As Figure 5 shown, the two-dimensional residual network sequentially includes a convolutional block as the input layer, a residual block composed of multiple convolutional blocks and a cross-layer network connection (where Figure 5 taking the residual block composed of 3 convolutional blocks and a cross-layer network connection as an example), and a two-dimensional convolutional layer as the output layer. Each convolutional block includes a two-dimensional convolutional layer and an activation function.
[0081] Among them, the residual blocks in the two-dimensional residual network are multiple consecutive ( Figure 5 taking 6 consecutive ones as an example).
[0082] As Figure 5 shown, the human silhouette sequence is input into the convolutional block as the input layer. After being convolved by the two-dimensional convolutional kernel in this convolutional block, it is input into the activation function to obtain the output data of the activation function; the output data of the activation function is input into the first residual block for feature processing, and then the output data of the first residual block is input into the second residual block for feature processing, and so on, through the feature processing of multiple residual blocks. The output data of the last residual block is input into the convolutional block as the output layer to obtain the output data of this convolutional block, that is, the silhouette spatial domain features are obtained.
[0083] In a possible implementation, a silhouette spatial domain feature extraction network is used to extract features from a sequence of human silhouettes, obtaining silhouette spatial domain features, which can be expressed by the following formula:
[0084]
[0085] Among them, represents the sequence of human silhouettes, s sil represents the silhouette spatial domain features, T represents the number of frames in the sequence of human silhouettes, C represents the number of feature channels of the human silhouettes in the sequence of human silhouettes, N1 represents the number of feature channels of the silhouette spatial domain features, H1 represents the image length of the human silhouette, H1↓ represents the feature map length of the silhouette spatial domain features, W represents the image width of the human silhouette, and W↓ represents the feature map width of the silhouette spatial domain features. R represents the silhouette spatial domain feature extraction network.
[0086] S404. Input the silhouette spatial domain features into a pooling network, and perform feature pooling on the silhouette spatial domain features through the pooling network to obtain human silhouette features.
[0087] In this embodiment, the silhouette spatial domain features output by the silhouette spatial domain feature extraction network are input into the pooling network. In the pooling network, feature pooling is performed on the silhouette spatial domain features. Among them, the process of feature pooling is a process of feature aggregation and representation. Finally, human silhouette features are obtained.
[0088] In a possible implementation, as Figure 3 shown, the pooling network includes a temporal pooling (TP) network and a horizontal pyramid pooling (HPP) network, so as to use the temporal pooling network and the horizontal pyramid network to aggregate and represent the temporal features of the human silhouette features, so that the final human silhouette features are obtained through the extraction of spatial domain features and temporal features, improving the accuracy of the human silhouette features, and further improving the accuracy of personnel identity recognition.
[0089] Based on the fact that the pooling network includes a temporal pooling network and a horizontal pyramid pooling network, a possible implementation of S404 includes: inputting the silhouette spatial domain features into the temporal pooling network, performing max pooling on the silhouette spatial domain features in the temporal dimension through the temporal pooling network to obtain preliminary pooled features; inputting the preliminary pooled features into the horizontal pyramid pooling network, and in the horizontal pyramid network, performing multi-scale partitioning, average pooling (AvgPooling), max pooling, and pooling feature merging on the preliminary pooled features in the spatial dimension to obtain human silhouette features.
[0090] In this implementation manner, the process of feature pooling for the silhouette spatial domain features through the time-domain pooling network and the horizontal pyramid network can be expressed by the following formula:
[0091]
[0092] Among them, TP represents time-domain pooling, which mainly performs max pooling on the input silhouette spatial domain feature s sil in the time sequence dimension; HPP represents horizontal pyramid pooling, and t sil represents the silhouette feature obtained after passing through the HPP network. The processing process of HPP mainly includes: dividing the input initial pooling feature in the spatial dimension into M (M>1) different scales to obtain multiple divided units; for each divided unit, performing average pooling and max pooling in the spatial dimension; then, summing the results obtained by average pooling and max pooling to obtain the eigenvalue of this divided unit; finally, combining the eigenvalues of each divided unit, that is where K m represents the number of divided units obtained after the m-th division, and P is the feature dimension of the eigenvalue obtained after combination.
[0093] S405. Input the human key point sequence into the time-domain feature extraction network in the key point feature extraction network, and through the time-domain feature extraction network, extract the features of the human key point sequence to obtain the human key point features of the person.
[0094] In this embodiment, the human key point sequence is input into the time-domain feature extraction network in the key point feature extraction network, and the time-domain feature extraction network is used to extract the features in the time sequence dimension of the human key point sequence to obtain the human key point features
[0095] Among them, J represents the number of key points in each group of human key points, 2 represents the two-dimensional Euclidean space coordinate points, N2 represents the number of feature channels after passing through the time-domain feature extraction network, and T2↓ represents downsampling in the time sequence dimension after passing through the time-domain feature extraction network.
[0096] In an example, the network structure of the time-domain feature extraction network can refer to the network structure of the spatial domain feature extraction network and adopt a residual network. The difference is that the convolutional layer in the time-domain feature extraction network is a one-dimensional convolutional layer, so the time-domain feature extraction network can be a one-dimensional residual network. Modify the two-dimensional convolutional layer in the Figure 5 shown two-dimensional residual network to a one-dimensional convolutional layer, which is a one-dimensional residual network.
[0097] S406. Through the multi-modal feature mixing network in the identity feature extraction model, fuse the human silhouette feature and the human key point feature to obtain the identity feature of the person.
[0098] S407. Identify the identity of the person according to the identity feature.
[0099] Among them, the implementation principles and technical effects of S406 to S407 can refer to the foregoing embodiments and will not be elaborated here.
[0100] In the embodiments of the present disclosure, from the captured video of the robot, a human silhouette sequence and a human key point sequence reflecting the characteristics of human gait are extracted; in the silhouette feature extraction network of the identity feature extraction model, through the spatial domain feature extraction network and the pooling network, the human silhouette sequence is subjected to feature extraction to improve the extraction accuracy of the human silhouette feature; in the silhouette feature extraction network of the key point feature extraction model, through the time domain feature extraction network, the human key point sequence is subjected to feature extraction to improve the extraction accuracy of the human key point feature; through the multi-modal feature mixing network, the human silhouette sequence and the human key point sequence are subjected to feature extraction and feature fusion to obtain the identity feature, improving the diversity and accuracy of the identity feature. Finally, the identity of the person is recognized based on the identity feature. Thus, the convenience and accuracy of identity recognition are improved through gait monitoring on the robot.
[0101] In some embodiments, as Figure 3 shown, the multi-modal feature mixing network includes a convolutional network corresponding to the human silhouette feature, a convolutional network corresponding to the human key point feature, and an attention layer. Based on this, through the multi-modal feature mixing network in the identity feature extraction model, the human silhouette feature and the human key point feature are fused to obtain the identity feature of the person, including: in the multi-modal feature mixing network, through the convolutional network corresponding to the human silhouette feature, the convolutional network corresponding to the human key point feature, and the attention layer, the human silhouette feature and the human key point feature are fused to obtain the identity feature. Thus, the fusion effect of fusing the human silhouette feature and the human key point feature is improved by using the attention mechanism.
[0102] Among them, as Figure 3 shown, the human silhouette feature can correspond to one convolutional network, and the human key point feature can correspond to two convolutional networks. The convolutional network is used to calculate the statistics (such as mean, variance) of the feature by performing spatial transformation on the feature, and obtain the attention value (attention value) corresponding to the feature. Then, based on the attention value and the attention mechanism in the attention layer, feature fusion is realized.
[0103] In this embodiment, as Figure 3As shown, the human key point features can be respectively input into two corresponding convolutional networks (such as a convolutional layer with a convolution kernel size of 1*1). Through these two convolutional networks, the human key point features are respectively mapped to the first space (such as the g space) and the second space (such as the h space), and the statistic g of the human key points in the first space is calculated. pose and the statistic h of the human key points in the second space. pose The human silhouette feature can be input into its corresponding convolutional network (such as a convolutional layer with a convolution kernel size of 1x1). Through this convolutional layer, the human silhouette feature is mapped into the third space, and the statistic f of the human key points in the third space is calculated. sil . Then, in the attention layer, based on the statistic g of the human key point features in the first space pose and the statistic f of the human silhouette features in the third space sil , the attention value is determined; then the statistic h of the human key point features in the second space pose is weighted with the attention value to achieve the feature fusion of the human silhouette feature and the human key point feature. Among them, the first space, the second space, and the third space correspond to local statistics such as mean and variance, for example.
[0104] In some embodiments, as Figure 3 shown, the multi-modal feature mixing network further includes a fully connected layer. After the human silhouette feature and the human key point feature are fused, a fused feature is obtained. The fused feature is input into the fully connected layer to obtain the identity feature output by the fully connected layer.
[0105] In one example, the process of fusing the human silhouette feature and the human key point feature by the multi-modal feature mixing network to obtain the fused feature can be expressed by the following formula:
[0106] f sil = W f t sil
[0107] g pose = W g t pose
[0108] h pose = W h t pose
[0109]
[0110] Among them, W f is the network parameter of the convolutional network corresponding to the human silhouette feature, and W g are the network parameters of the two convolutional networks respectively corresponding to the human key point features. t spis a mixed feature, d f is f sil in the dimension of the feature channel. As an example, in f sil and g pose with a mean of 0 and a variance of 1, B = f sil T g pose also has a mean of 0 and a variance of std. When std increases, the variance of the elements in B also increases. To avoid the distribution of B from becoming too steep, by dividing by the variance of B becomes 1 again, improving the stability of B and thus the stability of the gradient during model training.
[0111] In some embodiments, in the recognition network, the KD-Tree algorithm can be used and with the help of the Euclidean distance metric, in the identity feature library, calculate the distance between the identity feature of the service object and the identity feature of the person in the captured video. If the distance is less than the distance threshold, it can be determined that the person is the service object (i.e., the registered user) of the robot. Thus, the accuracy of person recognition is improved.
[0112] In some embodiments, as Figure 3 shown, the identity feature extraction model may further include a classification network, which plays a role in the training process of the identity feature extraction model. During the model training process, input the identity feature into the classification network to obtain the predicted identity label of the person; based on the difference between the predicted identity label of the person and the true identity label of the person, determine the first loss value; based on the first loss value, adjust the model parameters of the identity feature extraction model to achieve model training.
[0113] In one example, the classification network is a batch normalization layer to improve the classification effect through the batch normalization layer.
[0114] In some embodiments, the loss value for training the identity feature extraction model further includes a second loss value. Based on the identity feature, the triplet loss function can be used to determine the second loss value. Based on the first loss value and the second loss value, adjust the model parameters of the identity feature extraction model to achieve model training. Among them, the triplet loss can aggregate the feature distances between different categories in the feature space, which is beneficial to improving the training effect of the identity feature model.
[0115] Figure 6 is the flowchart of the gait monitoring method based on the robot provided by the embodiments of the present disclosure Figure 3 . As Figure 6 shown, the gait monitoring method of the robot includes:
[0116] S601. Obtain the captured video of the robot.
[0117] S602. Perform image processing on the video frames of the captured video to obtain a sequence of human silhouettes and a sequence of human key points of the people in the captured video.
[0118] S603. Extract features from the sequence of human silhouettes through the silhouette feature extraction network in the identity feature extraction model to obtain the human silhouette features of the person.
[0119] S604. Extract features from the sequence of human key points through the key point feature extraction network in the identity feature extraction model to obtain the human key point features of the person.
[0120] S605. Perform feature fusion on the human silhouette features and the human key point features through the multi-modal feature mixing network in the identity feature extraction model to obtain the identity features of the person.
[0121] S606. Identify the identity of the person based on the identity features.
[0122] Among them, the implementation principles and technical effects of S601 to S605 can be referred to the foregoing embodiments and will not be elaborated herein.
[0123] S607. If the identity recognition result is that the person belongs to the service object, then identify the sequence of human key points through the health monitoring network to obtain the gait health status of the person.
[0124] In this embodiment, when the identity recognition result of the person is that the person belongs to the service object of the robot, that is, belongs to the authorized user of the robot, since the sequence of human key points can reflect the human gait clearly and in detail, the sequence of human key points can be identified through the health monitoring network to obtain the gait health status of the person. Among them, the input data of the health monitoring network is the sequence of human key points, and the output data is the gait health status of the person. The health monitoring network can be a deep neural network.
[0125] Thus, health monitoring of the person based on the person's gait on the robot is performed, improving the intelligence level of the robot and further improving the user experience.
[0126] In some embodiments, as Figure 7 ( Figure 7 is the structural schematic diagram of the health monitoring network provided by the embodiment of the present application), the health monitoring network includes a feature encoding network and a gait health recognition network. Based on this, as Figure 6 shown, a possible implementation manner of S607 includes:
[0127] S6071. If the identity recognition result of the person is that the person belongs to the service object, then perform feature encoding on the sequence of human key points through the feature encoding network in the health monitoring network to obtain encoded features.
[0128] Among them, the feature encoding network is a multi-layer convolutional network.
[0129] In this embodiment, if the identity recognition result of a person indicates that the person belongs to the service object, the human key point sequence of the person can be input into the feature encoding network of the health monitoring network. In the feature encoding network, the human key point sequence is feature-encoded to obtain encoded features.
[0130] In a possible implementation manner, as Figure 7 shown, the human key point sequence includes a facial key point sequence and a body key point sequence, and the feature encoding network in the gait health monitoring network includes a facial encoding network and a gait encoding network. Thus, by combining facial key points and body key points, the accuracy of the robot's gait health monitoring of a person through the gait health monitoring network is improved.
[0131] Among them, the body key point sequence may include body key points and limb key points.
[0132] Based on the fact that the human key point sequence includes a facial key point sequence and a body key point sequence, and the feature encoding network in the gait health monitoring network includes a facial encoding network and a gait encoding network, further, S6071 may include: If the identity recognition result of a person indicates that the person belongs to the service object, the facial key point sequence is feature-encoded through the facial encoding network to obtain facial features, and the body key point sequence is feature-encoded through the gait encoding network to obtain gait features.
[0133] As Figure 7 shown, the facial key point sequence is input into the facial encoding network to obtain the facial features output by the facial encoding network, and the body key point sequence is input into the gait encoding network to obtain the gait features output by the gait encoding network.
[0134] As an example, after obtaining the facial key points based on the captured video the facial encoding network (i.e., the facial key point encoder) can be used to analyze the temporal actions of the facial key points to obtain facial features (i.e., facial key point features) Among them, J face is the number of facial key points, and N face is the number of feature channels of the facial features.
[0135] As an example, after obtaining the body key points based on the captured video the gait encoding network (i.e., the gait key point encoder) can be used to analyze the temporal actions of the body key points to obtain gait features (i.e., pace key point features) Among them, J body is the number of body key points, and N bodyThe number of feature channels for gait features.
[0136] S6072. Identify the encoded features through the gait health recognition network in the health monitoring network to obtain the gait health status of the person.
[0137] Among them, the gait health recognition network includes a gait health detection network and a gait health classification network, and both the gait health detection network and the gait health classification network are multi-layer convolutional networks.
[0138] In this embodiment, the encoded features can be input into the gait health detection network of the gait health recognition network, and the gait health features can be obtained by extracting the features of the encoded features through the gait health detection network; then, the gait health features can be input into the gait health classification network to obtain the gait health of the person.
[0139] In a possible implementation, as Figure 7 shown, the gait health recognition network includes at least one of the following: a step recognition network, an emotion recognition network, and a leg type recognition network. Among them, the step recognition network can be used to identify the step status of the person, the emotion recognition network can be used to identify the emotion type of the person, and the leg type recognition network can be used to identify the leg type of the person. Based on this, the gait health status of the person includes at least one of the following: the step status of the person, the emotion type of the person, and the leg type of the person. Thus, the gait health of the person is recognized from one or more aspects, improving the accuracy of gait health recognition of the person.
[0140] Based on the gait health recognition network including at least one of the following: a step recognition network, an emotion recognition network, and a leg type recognition network, S6072 can include at least one of the following implementation methods:
[0141] Method 1: Identify the gait features through the step recognition network to obtain the step status of the person.
[0142] Method 2: Identify the facial features and gait features through the emotion recognition network to obtain the emotion type of the person. Thus, the accuracy of emotion type recognition of the person is improved by using the characteristics that both facial key points and gait key points reflect emotions.
[0143] Method 3: Identify the gait features through the leg type recognition network to obtain the leg type of the person.
[0144] Among them, as Figure 7 shown, the step recognition network can include a step health detection network and a step health classification network; the emotion recognition network can include an emotion health detection network and an emotion health classification network; the leg type recognition network can include a leg type health detection network and a leg type health classification network.
[0145] In Method 1, as Figure 7As shown, gait features can be first input into a gait health monitoring network for gait feature extraction, and then the output data of the gait health monitoring network can be input into a gait health classification network for classification to obtain the gait condition of the person.
[0146] In one example, as Figure 7 shown, the gait condition may include a uniform gait and / or a non-uniform gait. Thus, whether the person's gait is uniform can be identified through the gait recognition network, and the gait health condition of the person can be reflected from the perspective of whether the gait is uniform.
[0147] In one example, the gait health monitoring network mainly includes a fully connected layer. After receiving the encoded features, the gait health monitoring network can map the encoded features to another local space to extract gait features from the encoded features. During model training, the gait loss can be calculated using the predicted label obtained from the gait health classification network (i.e., predicting whether the person's gait is uniform) and the true label (i.e., whether the person's gait is actually uniform). Using the gait loss, the network parameters of the gait health recognition network and the feature encoding network can be adjusted so that they can determine whether the user's gait is uniform during model application.
[0148] In Method 2, as Figure 7 shown, facial features and gait features can be first input into an emotion health monitoring network for gait feature extraction. During feature extraction, the emotion-related features in the facial features and the emotion-related features in the gait features can be feature fused, for example, they can be fused by a weighted method to obtain the output data of the emotion health monitoring network; then the output data of the emotion health monitoring network can be input into an emotion health classification network for classification to obtain the emotion type of the person.
[0149] In one example, during the process of feature fusing the emotion-related features in the facial features and the emotion-related features in the gait features, it may include: using a convolutional network corresponding to the facial features to perform feature processing on the emotion-related features in the facial features; using a convolutional network corresponding to the gait features to perform feature processing on the emotion-related features in the gait features; weighting the output data of the convolutional network corresponding to the facial features and the output data of the convolutional network corresponding to the gait features to obtain a mixed feature, and then mapping the mixed feature to other feature spaces through a fully connected layer so that the mixed feature is beneficial to subsequent emotion classification based on the mixed feature.
[0150] Among them, the convolutional network corresponding to the facial features and the convolutional network corresponding to the gait features can use a fully connected layer, and can also include a batch normalization layer and an activation function, etc.
[0151] In one example, as Figure 7As shown, the preset emotion types may include one of the following: happy, natural, sad, angry. Therefore, the emotion type of the person can be identified as one of the preset emotion types.
[0152] In one example, when the person is at a long distance, the facial information may be incomplete or unsound. Therefore, when the facial information of the person is incomplete, the emotion type of the person can be determined according to the gait characteristics of the person. When the facial information of the person is complete, the gait characteristics of the person can be used to perform an initial classification of the emotion of the person to obtain the range of the emotion type of the person; then, based on the facial characteristics of the person, the emotion of the person is classified within the range of the emotion type of the person to obtain the emotion type of the person. Thus, it adapts to different scenarios, relies on different data for emotion recognition, ensures the recognition of the emotional health of the person in different scenarios, and ensures a certain degree of accuracy.
[0153] In the process of determining the emotion type of the person according to the gait characteristics of the person, the gait characteristics can be recognized for emotion through a gait emotion recognition network. The gait emotion recognition network may include multiple convolutional blocks, and each convolutional block may include a convolutional layer and an activation function. The gait characteristics are used to extract the emotion-related characteristics in the gait through three groups of convolutional blocks. Further, each convolutional block may include a one-dimensional convolutional layer and a leaky rectified linear units (Leaky ReLU) activation function.
[0154] In the process of classifying the emotion of the person within the range of the emotion type of the person based on the facial characteristics of the person, a facial emotion recognition network can be used to recognize the emotion of the facial characteristics to obtain the emotion-related characteristics in the facial characteristics. Among them, the facial emotion recognition network may be composed of multiple fully connected layers (such as 3 fully connected layers) to extract the emotion-related characteristics in the facial characteristics through these fully connected layers to improve the accuracy of feature extraction.
[0155] In Method 3, as Figure 7 shown, the gait characteristics can be input into a leg health monitoring network for leg feature extraction first, and then the output data of the leg health monitoring network is input into a leg health classification network for classification to obtain the leg type of the person.
[0156] As Figure 7 shown, the preset leg types may include at least one of the following: O-shaped legs, X-shaped legs, and normal leg types. Therefore, based on the leg health monitoring network and the leg health classification network, the leg type of the person can be identified as one of the preset leg types.
[0157] In one example, Figure 8 This is a schematic diagram of the data processing flow of the leg health monitoring network provided by the embodiment of the present application. As Figure 8As shown, in the leg type health monitoring network, gait features can be first input into a convolutional network (such as a 1x1 convolutional network), and the convolutional network is used to map the gait features to g body . Taking the 1x1 convolutional network as an example, this mapping process can be expressed as:
[0158] where J represents the number of body key points in each group of human key points in the key point sequence.
[0159] Meanwhile, a fully connected layer can be used to perform feature analysis on the gait features on the feature channels, and the output value of this fully connected layer is calculated through the softmax function to obtain the attention value in the leg type health monitoring network. After that, this attention value can be multiplied by the mapped feature g of the gait features body , and then the fully connected layer is used to process the multiplied features to obtain the leg type-related features in the gait features, that is, the leg type features.
[0160] After that, the leg type features are input into the leg type health classification network to obtain the leg type of the person.
[0161] Therefore, through the above method, one or more gait-based health monitoring is provided for the service object of the robot, effectively improving the accuracy and convenience of health monitoring, effectively improving the intelligence level of the robot, and further effectively improving the user experience of the service object.
[0162] In some embodiments, after obtaining the gait health status of a person, if the gait health status of the person meets the reminder condition, the robot can be controlled to perform corresponding reminder operations according to the gait health status of the person. Thus, the gait health of the person is timely reminded, improving the user experience.
[0163] In one example, the reminder condition may include at least one of the following: uneven steps, sad mood, angry mood, abnormal leg type (such as O-shaped legs, X-shaped legs).
[0164] Thus, health monitoring and reminder in one or more aspects of steps, mood, and leg type can be provided for the user, improving the intelligence level of the robot and the user experience.
[0165] In one example, controlling the robot to perform corresponding reminder operations according to the gait health status of the person may include at least one of the following;
[0166] (1) Based on the uneven pace of the person, the robot stores the uneven gait of the person and reminds the user of the uneven pace through voice, messages, etc. The robot can also display a health training video for adjusting the pace on the display screen and show the user's pace recovery situation in each time period, such as showing the user's pace recovery situation weekly or monthly.
[0167] (2) Based on the person's sad or angry mood, the robot stores the mood type of the person and interacts with the user through voice, messages, actions, etc. to improve the person's mood. The robot can also push entertainment content to the user, such as entertainment videos and music, etc., and show the user's mood adjustment situation in each time period.
[0168] (3) Based on the abnormal leg shape of the person, the robot stores the leg shape of the person and interacts with the user through voice, images, actions, etc. to remind the user to improve the leg shape. The robot can also push leg shape rehabilitation content and show the user's leg shape adjustment situation in each time period.
[0169] Corresponding to the robot-based gait monitoring method in the above embodiments, Figure 9 is the structural block diagram of the robot-based gait monitoring device provided by the embodiments of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown. Referring to Figure 9 , the robot-based gait monitoring device includes: a video acquisition unit 901, a video processing unit 902, a silhouette feature extraction unit 903, a key point feature extraction unit 904, a feature fusion unit 905, and an identity recognition unit 906.
[0170] The video acquisition unit 901 is used to acquire the shooting video of the robot;
[0171] The video processing unit 902 is used to perform image processing on the video frames of the shooting video to obtain the human silhouette sequence and the human key point sequence of the person in the shooting video;
[0172] The silhouette feature extraction unit 903 is used to extract features from the human silhouette sequence through the silhouette feature extraction network in the identity feature extraction model to obtain the human silhouette features of the person;
[0173] The key point feature extraction unit 904 is used to extract features from the human key point sequence through the key point feature extraction network in the identity feature extraction model to obtain the human key point features of the person;
[0174] The feature fusion unit 905 is used to perform feature fusion on the human silhouette features and the human key point features through the multi-modal feature mixing network in the identity feature extraction model to obtain the identity features of the person;
[0175] An identity recognition unit 906 for recognizing the identity of a person according to identity features.
[0176] In some embodiments, the silhouette feature extraction network includes a spatial domain feature extraction network and a pooling network. The silhouette feature extraction unit 903 is specifically configured to: input a sequence of human silhouettes into the spatial domain feature extraction network, extract features from the sequence of human silhouettes in the spatial domain feature extraction network to obtain silhouette spatial domain features; input the silhouette spatial domain features into the pooling network, and perform feature pooling on the silhouette spatial domain features through the pooling network to obtain human silhouette features.
[0177] In some embodiments, the pooling network includes a temporal pooling network and a horizontal pyramid pooling network. In the process of inputting the silhouette spatial domain features into the pooling network and performing feature pooling on the silhouette spatial domain features through the pooling network to obtain human silhouette features, the silhouette feature extraction unit 903 is specifically configured to: input the silhouette spatial domain features into the temporal pooling network, perform max pooling on the silhouette spatial domain features in the temporal dimension through the temporal pooling network to obtain preliminary pooled features; input the preliminary pooled features into the horizontal pyramid pooling network, and in the horizontal pyramid network, perform multi-scale partitioning, average pooling, max pooling, and pooled feature merging on the preliminary pooled features in the spatial dimension to obtain human silhouette features.
[0178] In some embodiments, the multi-modal feature mixing network includes convolutional layers corresponding to human silhouette features and human key point features respectively. The feature fusion unit 905 is specifically configured to: in the multi-modal feature mixing network, perform feature fusion on the human silhouette features and human key point features through the convolutional layer corresponding to the human silhouette features, the convolutional layer corresponding to the human key points, and an attention mechanism to obtain identity features.
[0179] In some embodiments, the gait monitoring device based on a robot further includes: a health monitoring unit 907 for, if the identity recognition result indicates that the person belongs to a service object, recognizing the sequence of human key points through a health monitoring network to obtain the gait health condition of the person.
[0180] In some embodiments, the health monitoring network includes a feature encoding network and a gait health recognition network. The gait health recognition network includes at least one of the following: a step recognition network, an emotion recognition network, and a leg type recognition network. The health monitoring unit 907 is specifically configured to: if the identity recognition result indicates that the person belongs to a service object, perform feature encoding on the sequence of human key points through the feature encoding network to obtain encoded features; recognize the encoded features through the gait health recognition network to obtain the gait health condition, and the gait health condition includes at least one of the following: the step condition of the person, the emotion type of the person, and the leg type of the person.
[0181] In some embodiments, the sequence of human key points includes the sequence of facial key points and the sequence of body key points, and the feature encoding network includes the facial encoding network and the gait encoding network. In the process of performing feature encoding on the sequence of human key points through the feature encoding network to obtain encoded features if the identity recognition result indicates that the person belongs to the service object, the health monitoring unit 907 is specifically configured to: if the identity recognition result indicates that the person belongs to the service object, perform feature encoding on the sequence of facial key points through the facial encoding network to obtain facial features, and perform feature encoding on the sequence of body key points through the gait encoding network to obtain gait features. In the process of identifying the encoded features through the gait health recognition network to obtain the gait health condition, the health monitoring unit 907 is specifically configured to: identify the gait features through the gait recognition network to obtain the gait condition; identify the facial features and the gait features through the emotion recognition network to obtain the emotion type; identify the gait features through the leg type recognition network to obtain the leg type.
[0182] The gait monitoring device based on a robot provided in this embodiment can be used to implement the technical solutions of the embodiments of the above-mentioned gait monitoring method based on a robot. The implementation principle and technical effects are similar, and will not be elaborated here.
[0183] Reference Figure 10 , which shows a schematic structural diagram of an electronic device 1000 suitable for implementing the embodiments of the present disclosure. The electronic device 1000 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs and desktop computers. Figure 10 The electronic device shown is only an example and should not impose any limitations on the functions and usage scopes of the embodiments of the present disclosure.
[0184] Such as Figure 10As shown, the electronic device 1000 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1001, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are also stored. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.
[0185] Generally, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 10 an electronic device 1000 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively.
[0186] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0187] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0188] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device.
[0189] The above-mentioned computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above-mentioned embodiments.
[0190] Computer program code for performing the operations of this disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0191] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0192] The units involved in the embodiments described in this disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the acquisition unit may also be described as "the unit for acquiring the page image and page description text of the web page to be detected".
[0193] The functions described above in this document may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and so on.
[0194] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0195] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0196] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0197] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A robot-based gait monitoring method, characterized in that, Including: Obtaining the captured video of the robot; Performing image processing on the video frames of the captured video to obtain a sequence of human silhouettes and a sequence of human key points of the people in the captured video, where the sequence of human key points includes a sequence of facial key points and a sequence of body key points; Extracting features from the sequence of human silhouettes through a silhouette feature extraction network in the identity feature extraction model to obtain the human silhouette features of the people; Extracting features from the sequence of human key points through a key point feature extraction network in the identity feature extraction model to obtain the human key point features of the people; Performing feature fusion on the human silhouette features and the human key point features through a multi-modal feature mixing network in the identity feature extraction model to obtain the identity features of the people; Performing identity recognition on the people according to the identity features; If the identity recognition result is that the people belong to the service object, then identifying the sequence of human key points through a health monitoring network to obtain the gait health status of the people, where the health monitoring network includes a feature encoding network and a gait health recognition network, the feature encoding network includes a facial encoding network and a gait encoding network, and the gait health recognition network includes a pace recognition network, an emotion recognition network, and a leg type recognition network; If the identity recognition result is that the people belong to the service object, then identifying the sequence of human key points through a health monitoring network to obtain the gait health status of the people, including: If the identity recognition result is that the people belong to the service object, then performing feature encoding on the sequence of facial key points through the facial encoding network to obtain facial features, and performing feature encoding on the sequence of body key points through the gait encoding network to obtain gait features; Identifying the gait features through the pace recognition network to obtain the pace status of the people; Identifying the facial features and the gait features through the emotion recognition network to obtain the emotion type of the people; Identifying the gait features through the leg type recognition network to obtain the leg type of the people.
2. The robot-based gait monitoring method according to claim 1, wherein, The silhouette feature extraction network includes a spatial domain feature extraction network and a pooling network. Extracting features from the sequence of human silhouettes through the silhouette feature extraction network in the identity feature extraction model to obtain the human silhouette features of the people includes: Inputting the sequence of human silhouettes into the spatial domain feature extraction network, and extracting features from the sequence of human silhouettes in the spatial domain feature extraction network to obtain silhouette spatial domain features; Inputting the silhouette spatial domain features into the pooling network, and performing feature pooling on the silhouette spatial domain features through the pooling network to obtain the human silhouette features.
3. The gait monitoring method based on a robot according to claim 2, wherein, The pooling network includes a temporal pooling network and a horizontal pyramid pooling network. Inputting the silhouette spatial domain features into the pooling network, and performing feature pooling on the silhouette spatial domain features through the pooling network to obtain the human silhouette features includes: Input the silhouette airspace feature into the temporal pooling network, and perform max pooling on the silhouette airspace feature in the temporal dimension through the temporal pooling network to obtain a preliminary pooling feature; Input the preliminary pooling feature into the horizontal pyramid pooling network. In the horizontal pyramid network, perform multi-scale partitioning, average pooling, max pooling, and pooling feature merging on the preliminary pooling feature in the spatial dimension to obtain the human silhouette feature.
4. The robot-based gait monitoring method according to any one of claims 1-3, characterized in that, The multi-modal feature mixing network includes a convolutional layer corresponding to the human silhouette feature, a convolutional layer corresponding to the human key point feature, and an attention layer. Through the multi-modal feature mixing network in the identity feature extraction model, perform feature fusion on the human silhouette feature and the human key point feature to obtain the identity feature of the person, including: In the multi-modal feature mixing network, perform feature fusion on the human silhouette feature and the human key point feature through the convolutional network corresponding to the human silhouette feature, the convolutional network corresponding to the human key point feature, and the attention layer to obtain the identity feature.
5. A robot-based gait monitoring device, characterized in that, Including: A video acquisition unit for acquiring the captured video of the robot; A video processing unit for performing image processing on the video frames of the captured video to obtain a human silhouette sequence and a human key point sequence of the person in the captured video. The human key point sequence includes a facial key point sequence and a body key point sequence; A silhouette feature extraction unit for extracting features from the human silhouette sequence through the silhouette feature extraction network in the identity feature extraction model to obtain the human silhouette feature of the person; A key point feature extraction unit for extracting features from the human key point sequence through the key point feature extraction network in the identity feature extraction model to obtain the human key point feature of the person; A feature fusion unit for performing feature fusion on the human silhouette feature and the human key point feature through the multi-modal feature mixing network in the identity feature extraction model to obtain the identity feature of the person; An identity recognition unit for performing identity recognition on the person according to the identity feature; A gait condition obtaining unit for, if the identity recognition result is that the person belongs to the service object, performing recognition on the human key point sequence through the health monitoring network to obtain the gait health condition of the person. The health monitoring network includes a feature encoding network and a gait health recognition network. The feature encoding network includes a facial encoding network and a gait encoding network. The gait health recognition network includes a pace recognition network, an emotion recognition network, and a leg type recognition network; When the gait condition obtaining unit performs recognition on the human key point sequence through the health monitoring network to obtain the gait health condition of the person, it includes: If the identity recognition result is that the person belongs to the service object, perform feature encoding on the facial key point sequence through the facial encoding network to obtain a facial feature, and perform feature encoding on the body key point sequence through the gait encoding network to obtain a gait feature; The gait characteristics are identified through the gait recognition network to obtain the gait condition of the person; The facial features and the gait characteristics are identified through the emotion recognition network to obtain the emotion type of the person; The gait characteristics are identified through the leg type recognition network to obtain the leg type of the person.
6. An electronic device, characterized in that, Comprising: At least one processor and a memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the robot-based gait monitoring method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the processor executes the computer-executable instructions, the robot-based gait monitoring method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Gait recognition method and device
CN114140883A
Image recognition method and apparatus, and terminal and storage medium
WO2019174439A1