Action recognition method and device, vehicle control method and device

By identifying the skeletal key points and periodic features of the target object in the image frame, the problems of sensor wearing inconvenience and lighting effects are solved, and high-accuracy motion recognition under changing lighting conditions is achieved, which is suitable for intelligent interaction and intelligent control.

CN117058762BActive Publication Date: 2025-12-19BEIJING YINWO AUTOMOBILE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311035488.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-16
Publication Date
2025-12-19
Estimated Expiration
2043-08-16

AI Technical Summary

Technical Problem

In existing technologies, sensor recognition requires the target object to wear a sensor, which is inconvenient, while image recognition is easily affected by lighting, resulting in inaccurate motion recognition results.

Method used

By acquiring target image frames, the skeletal key points of the target object are determined. Combined with the skeletal key point features in historical image frames, the action of the target object is identified. The spatial position and periodic features of the skeletal key points are used for action recognition, avoiding the influence of external environmental factors such as lighting.

Benefits of technology

It improves the accuracy of motion recognition, enabling accurate identification of target object movements under external environmental factors such as changes in lighting, and is suitable for intelligent interaction and intelligent control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058762B_ABST
    Figure CN117058762B_ABST
Patent Text Reader

Abstract

The application provides a motion recognition method and device, and a vehicle control method and device, and relates to the technical field of computers. The motion recognition method comprises the following steps: obtaining a target image frame, the target image frame containing a target object; determining, based on the target image frame, a skeletal key point of the target object in the target image frame; determining, based on the skeletal key point of the target object in the target image frame, a first spatial position feature corresponding to the motion of the target object; determining a periodic feature corresponding to the motion of the target object, the periodic feature comprising a second spatial position feature corresponding to the motion of the target object in a historical image frame before the target image frame; and recognizing the motion of the target object based on the first spatial position feature and the periodic feature. The application determines the motion feature of the target object based on the skeletal key point, and recognizes the motion of the target object in combination with the motion feature of the target object in the historical image frame, which can greatly improve the recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a motion recognition method and device, and a vehicle control method and device. BACKGROUND

[0002] With the development of information technology, more and more application scenarios need to recognize the behavior of a target object (such as a human body) to realize intelligent interaction or intelligent control. In the related art, motion recognition is usually performed through sensor recognition or image recognition.

[0003] However, the sensor recognition method requires the target object to wear a corresponding sensor, which is inconvenient. The image recognition method is affected by factors such as light, resulting in inaccurate motion recognition results. SUMMARY

[0004] To solve the above technical problems, the present application is proposed. The embodiments of the present application provide a motion recognition method and device, and a vehicle control method and device.

[0005] In a first aspect, the embodiments of the present application provide a motion recognition method, comprising: obtaining a target image frame, the target image frame containing a target object; determining a skeletal key point of the target object in the target image frame based on the target image frame; determining a first spatial position feature corresponding to a motion of the target object based on the skeletal key point of the target object in the target image frame; determining a periodic feature corresponding to the motion of the target object, the periodic feature including a second spatial position feature corresponding to the motion of the target object in a historical image frame before the target image frame; and recognizing the motion of the target object based on the first spatial position feature and the periodic feature.

[0006] In a possible implementation, the determining of the periodic feature corresponding to the motion of the target object comprises: extracting a first intermediate feature based on the skeletal key point of the target object in the target image frame; extracting a second intermediate feature based on an attention mechanism and the skeletal key point of the target object in the target image frame, the second intermediate feature including a second spatial position feature extracted based on the attention mechanism and the skeletal key point of the target object in the historical image frame; and fusing the first intermediate feature and the second intermediate feature to obtain the periodic feature corresponding to the motion of the target object.

[0007] In a possible implementation, the recognizing of the motion of the target object based on the first spatial position feature and the periodic feature comprises: fusing the first spatial position feature and the periodic feature to obtain a motion feature corresponding to the target object; determining at least one motion feature sample; and recognizing the motion of the target object based on the motion feature corresponding to the target object and the at least one motion feature sample, the motion feature sample being determined based on a plurality of image frames corresponding to a complete motion.

[0008] In a possible implementation, determining the at least one action feature sample includes: obtaining an action video corresponding to each of the at least one action, each action video including a plurality of image frames corresponding to a complete action; determining a first spatial position feature and a periodic feature corresponding to each of the at least one action based on the action video corresponding to each of the at least one action; and fusing the first spatial position feature and the periodic feature corresponding to each of the at least one action to obtain an action feature sample corresponding to the action.

[0009] In a possible implementation, after the target image frame is obtained, the method further includes: determining identity information of the target object in the target image frame; and identifying the action of the target object based on the first spatial position feature and the periodic feature includes: determining the action of the target object from a plurality of preset actions corresponding to the identity information of the target object based on the first spatial position feature and the periodic feature.

[0010] In a second aspect, an embodiment of the present application provides a vehicle control method, including: determining an action of a target object in a target image based on the action recognition method mentioned in the first aspect or any possible implementation of the first aspect; determining a vehicle control instruction corresponding to the action of the target object; and controlling a target vehicle based on the vehicle control instruction.

[0011] In a third aspect, an embodiment of the present application provides an action recognition device, including: an obtaining module configured to obtain a target image frame, the target image frame including a target object; a first determining module configured to determine a skeletal key point of the target object in the target image frame based on the target image frame; a second determining module configured to determine a first spatial position feature corresponding to an action of the target object based on the skeletal key point of the target object in the target image frame; a third determining module configured to determine a periodic feature corresponding to the action of the target object, the periodic feature including a second spatial position feature corresponding to the action of the target object in a historical image frame before the target image frame; and an identifying module configured to identify the action of the target object based on the first spatial position feature and the periodic feature.

[0012] In a fourth aspect, an embodiment of the present application provides a vehicle control device, including: a first determining module configured to determine an action of a target object in a target image based on the action recognition method mentioned in the first aspect or any possible implementation of the first aspect; a second determining module configured to determine a vehicle control instruction corresponding to the action of the target object; and a control module configured to control a target vehicle based on the vehicle control instruction.

[0013] In a fifth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program for executing the method in the first aspect and the second aspect.

[0014] In a sixth aspect, an electronic device is provided, and the electronic device includes a processor, a memory for storing processor-executable instructions, and the processor is configured to execute the method of the first aspect and the second aspect.

[0015] The action recognition method provided in the embodiments of the present application can first acquire a target image frame containing a target object, then determine the skeletal key points of the target object in the target image frame based on the target image frame, and then determine the first spatial position feature corresponding to the action of the target object based on the skeletal key points of the target object in the target image frame, and determine the periodic feature corresponding to the action of the target object, and finally recognize the action of the target object based on the first spatial position feature and the periodic feature.

[0016] Since the skeletal key points are not color features of the image itself, they are not easily affected by external environmental factors such as light, and therefore, determining the action features of the target object based on the skeletal key points can greatly avoid the influence of external environmental factors on the recognition accuracy. In addition, since an action is usually a coherent behavior composed of multiple postures, the method of recognizing the action of the target object in combination with the action features of the target object in the historical image frames can greatly improve the recognition accuracy. BRIEF DESCRIPTION OF DRAWINGS

[0017] The above and other objects, features and advantages of the present application will become more apparent from the following detailed description of the application when taken in conjunction with the accompanying drawings. The drawings provided in the specification and the contents of the specification serve as part of the description of embodiments of the present application, and are used to explain the present application together with the embodiments of the present application, but do not constitute a limitation on the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1 Fig. 1 shows a schematic diagram of an LSTM network provided by an exemplary embodiment of the present application.

[0019] Figure 2 Fig. 2 shows a schematic diagram of an application scenario of an action recognition method provided by an exemplary embodiment of the present application.

[0020] Figure 3 Fig. 3 shows a flowchart of an action recognition method provided by an exemplary embodiment of the present application.

[0021] Figure 4 Fig. 4 shows a flowchart of determining a periodic feature provided by an exemplary embodiment of the present application.

[0022] Figure 5 Fig. 5 shows a structure diagram of a periodic feature extraction network provided by an exemplary embodiment of the present application.

[0023] Figure 6 Fig. 1 shows a flowchart of determining an action of a target object according to an example embodiment of the present application.

[0024] Figure 7 Fig. 2 shows a flowchart of determining at least one action feature sample according to an example embodiment of the present application.

[0025] Figure 8 Fig. 3 shows a flowchart of a vehicle control method according to an example embodiment of the present application.

[0026] Figure 9 Fig. 4 shows a structure diagram of an action recognition device according to an example embodiment of the present application.

[0027] Figure 10 Fig. 5 shows a structure diagram of a vehicle control device according to an example embodiment of the present application.

[0028] Figure 11 Fig. 6 shows a structure diagram of an electronic device according to an example embodiment of the present application. DETAILED DESCRIPTION

[0029] The technical solutions in the embodiments of the present application will be clearly and completely described in connection with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0030] SUMMARY

[0031] In the related art, the following methods are usually used for action recognition of a target object.

[0032] Method one is to wear a sensor on the target object, and then determine the action of the target object through the spatial position data collected by the sensor. However, this method has a high cost, and the target object needs to wear a sensor, which is relatively inconvenient.

[0033] Method two is to detect the action of the target object in the current image frame through an image detection model. However, light will affect the target image frame, thereby reducing the accuracy of the recognition result output by the image detection model. In addition, the image detection model usually only determines the action according to the features of the current image frame, while the action of the target object is continuous, and only the action features in the image frame at a certain moment cannot accurately recognize the action of the target object.

[0034] Method three, detecting the action of the target object through a long short-term memory (LSTM) network. As shown in Figure 1 The LSTM network contains features of historical image frames collected before the current image frame. However, as the image frames are collected, the LSTM network compresses the features extracted from the historical image frames, and thus the LSTM network is not good at extracting features of complete actions occurring in a period of time, and thus cannot accurately identify the action of the target object.

[0035] Therefore, the present application provides an action recognition method. The action recognition method can first acquire a target image frame containing a target object, then determine the skeletal key points of the target object in the target image frame based on the target image frame, and then determine the first spatial position features corresponding to the action of the target object based on the skeletal key points of the target object in the target image frame, and determine the periodic features corresponding to the action of the target object, and finally recognize the action of the target object based on the first spatial position features and the periodic features.

[0036] Since the skeletal key points are not color features of the image itself, they are not easily affected by external environmental factors such as light, and thus determining the action features of the target object based on the skeletal key points can greatly avoid the influence of external environmental factors on the recognition accuracy. In addition, since an action is usually a coherent behavior composed of multiple postures, the way of recognizing the action of the target object in combination with the action features of the target object in the historical image frames can greatly improve the recognition accuracy.

[0037] Exemplary scenarios

[0038] The action recognition method provided by the embodiments of the present application can be executed by an electronic device. The electronic device can be a terminal (user end), such as a smart phone, a tablet computer, a desktop computer, etc. or the electronic device can also be a server. In a possible application scenario, the action recognition method provided by the embodiments of the present application can be applied to a target application program in the terminal. In some other possible implementation manners, the action recognition method can be realized by a processor calling computer readable instructions stored in a memory.

[0039] Referring to Figure 2 FIG. 1 shows an application scenario diagram of the action recognition method provided by an exemplary embodiment of the present application. As shown in Figure 2As shown, the application embodiment provided application scenarios at least include: image acquisition module 21 and action detection module 22. Wherein, the image acquisition module 21 can include at least one image acquisition device, the image acquisition device is exemplarily a camera, the image acquisition module 21 is used to acquire the target image frame to be detected. The action detection module 22 is used to detect the action of the target object in the target image frame. Exemplarily, the target object can be a traffic police, and the action of the target object can be a hand gesture of the traffic police.

[0040] Exemplarily, in the actual application process, the image acquisition module 21 can send the target image frame to the action detection module 22 after obtaining the target image frame, and the action detection module 22 detects the action of the target object in the target image frame after receiving the target image frame.

[0041] In a possible application scenario, the application scenario can also include a processor 23, and the processor 23 is used to generate a corresponding control instruction based on the action of the target object. Exemplarily, in the actual application process, the action detection module 22 can send the action of the target object to the processor 23 after determining the action of the target object, so that the processor 23 generates a corresponding control instruction based on the action of the target object.

[0042] In a possible application scenario, the application scenario can also include a target device 24, and the processor 23 can control the target device 24 based on the control instruction. Exemplarily, the target device 24 can be a target vehicle, and the control instruction can be used to control the driving device (such as acceleration, deceleration, turning, etc.), display device (such as displaying prompt information), voice playing device (such as playing audio) and the like of the target vehicle.

[0043] Exemplary methods

[0044] Referring to Figure 3 As shown, the flowchart of the action recognition method provided by an exemplary embodiment of the application includes the following steps 301 to 305.

[0045] Step 301, obtaining a target image frame. Wherein, the target image frame contains a target object.

[0046] Step 302, determining the skeletal key points of the target object in the target image frame based on the target image frame.

[0047] Step 303, determining the first spatial position feature corresponding to the action of the target object based on the skeletal key points of the target object in the target image frame.

[0048] In step 304, a periodic feature corresponding to the action of the target object is determined. The periodic feature includes a second spatial position feature of the action of the target object in a historical image frame before the target image frame.

[0049] In step 305, the action of the target object is identified based on the first spatial position feature and the periodic feature.

[0050] The following will be described in detail with respect to step 301.

[0051] In some implementations, the target image frame can be acquired by a pre-set image acquisition device, which can be a camera. Alternatively, the image acquisition device can continuously acquire a plurality of image frames, and the target image frame can be any image frame containing the target object in the plurality of image frames.

[0052] It should be noted that the target object can be different in different application scenarios. For example, in a vehicle driving scenario, the target object can be a traffic police, and in an industrial operation scenario, the target object can be an operator.

[0053] It can be understood that if the image acquisition device continuously acquires a plurality of image frames, and the plurality of image frames include an image frame in which the target object does not exist, it is not necessary to determine the skeletal key points of the target object in the image frame in which the target object does not exist, and it is also not necessary to determine the skeletal key points of other objects in the plurality of image frames except the target object, so as to prevent interference with the detection of the skeletal key points of the target object mentioned in step 202. Therefore, in a possible implementation, after any image frame (including the target image frame) is acquired, it can be detected whether the target object is included in the image frame, and after it is detected that the target object exists in the image frame, the skeletal key points of the target object in the image frame are detected.

[0054] In a possible implementation, the target object in the image frame can be detected based on a target detection model, which can be a YOLOv5 model. In addition to determining whether the target object exists in the image frame, the target detection model can also output position information of the target object in the image frame (such as position information of a rectangular detection frame enclosing the target object), so as to accurately locate the target object based on the position information in step 302, and avoid interference of other objects when the skeletal key points of the target object are extracted.

[0055] In a possible implementation, the target detection model can be trained by the following steps: obtaining a sample image carrying label information; wherein the label information is used to indicate the position information of the target object in the sample image; inputting the sample image into the target detection model to be trained to obtain an object prediction result output by the target detection model to be trained; determining a first loss value based on the object prediction result and the label information, and adjusting the parameters of the target detection model based on the first loss value.

[0056] The step 302 is described below in detail.

[0057] The skeleton key points to be detected can be preset key points. For example, the target object is a human body, if the gesture action of the target object is to be recognized, the skeleton key points can be the key points of the upper body, and if the whole body action of the target object is to be recognized, the skeleton key points can be the key points of the whole body. In addition, when the skeleton key points of the target object are determined, the skeleton key points can be detected based on any human key point detection algorithm. The human key point detection algorithm can be an Alphapose algorithm, and the present application does not limit this.

[0058] After the position information of the target object in the target image frame is determined based on any method, the skeleton key points of the target object can be detected based on the position information of the target object. For example, the target region where the target object is located can be cut from the target image frame based on the position information of the target object to obtain a target region image, and the skeleton key points of the target object in the target region image are detected. In this way, the skeleton key points of the target object can be accurately detected, and the interference of other objects on the detection result can be avoided.

[0059] The step 303 is described below in detail.

[0060] The first spatial position feature can be obtained based on a deep learning network (hereinafter referred to as a spatial feature extraction network for the sake of distinction) having good extraction ability for spatial features. The first spatial position feature can represent the relationship of the plurality of skeleton key points in space. For example, the spatial feature extraction network can be a Spatial Temporal Graph Convolutional Networks (ST-GCN).

[0061] The step 304 is described below in detail.

[0062] Periodic features are used to characterize the features of skeletal keypoints in the time dimension, that is, the changes of skeletal keypoints over a period of time prior to the current time. In other words, they can be the features of the skeletal keypoints of the target object in historical image frames (i.e., second spatial location features). Historical image frames are image frames acquired before the target image frame. During the image acquisition process, historical image frames and target image frames can be multiple consecutive image frames.

[0063] In one possible implementation, when determining the periodic features corresponding to the action of the target object, the periodic features can be determined based on the skeletal key points of the target object in the target image frame. The periodic features also include the features of the skeletal key points in the target image frame; that is, the periodic features can include the features of the skeletal key points of the target object in the current target image frame, as well as the features of the skeletal key points of the target object in historical image frames (i.e., second spatial location features).

[0064] In some implementations, periodic features can be extracted based on periodic feature extraction networks. The following section will combine... Figure 4 and Figure 5 For example, illustrate how to determine the periodic characteristics corresponding to the actions of a target object.

[0065] Figure 4 The diagram shown is a flowchart illustrating the determination of periodic features according to an exemplary embodiment of this application. Figure 5 The diagram shown is a schematic diagram of the structure of a periodic feature extraction network provided in an exemplary embodiment of this application.

[0066] like Figure 4 As shown, to determine the periodic characteristics corresponding to the action of the target object, the following steps 401 to 403 can be performed.

[0067] Step 401: Extract the first intermediate feature based on the skeletal key points of the target object in the target image frame.

[0068] The first intermediate feature can be the feature corresponding to the skeletal key points of the target object in the current target image frame, or the first intermediate feature can also include the feature corresponding to the skeletal key points in some historical image frames (such as 3 historical image frames) before the target image frame. For example, the first intermediate feature is extracted according to the network structure of the LSTM network. The first intermediate feature is mainly used to characterize the features of the skeletal key points in the current or short term.

[0069] In a specific example, such as Figure 5 The first intermediate feature extraction module shown can extract the first intermediate feature based on the forget gate (f t ), Input gate (i t ), time unit ( and ), a storage unit (g t ), a hidden layer , etc., to extract the first intermediate feature. Wherein, t represents the time, and l represents the level.

[0070] In step 402, the second intermediate feature is extracted based on the attention mechanism and the skeleton key points of the target object in the target image frame.

[0071] The second intermediate feature includes the second spatial position feature extracted based on the attention mechanism and the skeleton key points of the target object in the historical image frame. Specifically, the attention mechanism is used to focus on the action features of the target object in the historical image frame. The attention mechanism can store the second spatial position feature corresponding to the historical image frame based on the storage unit, and generate the second intermediate feature based on the second spatial position feature. Alternatively, for any historical image frame, the extraction method of the second spatial position feature is the same as the method of extracting the second intermediate feature of the target image frame.

[0072] When generating the second intermediate feature based on the second spatial position feature, the third intermediate feature of the skeleton key points of the target object in the target image frame can be extracted first, and then the second intermediate feature is obtained based on the third intermediate feature and the stored second spatial position feature.

[0073] In a specific example, as shown in the second intermediate feature extraction module Figure 5 , the second intermediate feature can be extracted based on the forgetting gate (f t ), the input gate (i t ), the space-time memory unit (g and ), the storage unit (g t ), etc. Wherein, t represents the time, and l represents the level (the target image frame can be processed in multiple layers to obtain multiple second intermediate features). Specifically, for the feature of the l-1 level (i.e. the output feature), the feature can be multiplied with the parameters of the forgetting gate to obtain the first dot product result (i.e. the third intermediate feature of the l level), and the second spatial position feature stored in the storage unit can be multiplied with the parameters of the input gate to obtain the second dot product result. Finally, the first dot product result and the second dot product result are added to obtain the second intermediate feature of the l level.

[0074] After obtaining the second intermediate feature by any method, the second intermediate feature can be stored (such as stored in the storage unit) to generate the second intermediate feature of the image frame in which the action of the target object is identified after the target image frame.

[0075] Step 403: fusing the first intermediate feature and the second intermediate feature to obtain a periodic feature corresponding to the action of the target object.

[0076] It should be noted that the second intermediate feature can be the second intermediate feature of the last level or the second intermediate features of multiple levels.

[0077] By using this method, the features corresponding to the skeleton key points of the target object in the current target image frame and the features corresponding to the skeleton key points of the target object in the historical image frame can be accurately fused, so as to improve the accuracy of action recognition.

[0078] It can be understood that in the embodiments of the present application, the execution order of steps 303 and 304 is not limited, step 303 can be executed before or after step 304, or can be executed simultaneously with step 304.

[0079] The following will be described in detail with respect to step 305.

[0080] In some implementations, the first spatial position feature and the periodic feature can be input into a classifier, and the classifier outputs the action of the target object as probabilities of each preset action, and then determines the action of the target object based on the probabilities of each preset action. For example, the preset action with the highest probability can be taken as the action of the target object, or the preset action with a probability greater than a corresponding preset threshold can be taken as the action of the target object. Wherein, the preset threshold corresponding to each preset action can be the same or different. In addition, the classifier can be a normalized exponential function (softmax) classifier.

[0081] In one possible implementation, when identifying the action of the target object based on the first spatial position feature and the periodic feature, as shown in FIG. 6A, the following steps 601 to 603 can be performed. Figure 6

[0082] Step 601: fuse the first spatial position feature and the periodic feature to obtain an action feature corresponding to the target object.

[0083] For example, the first spatial position feature and the periodic feature can be fused to obtain an action feature corresponding to the target object according to the respective weights of the first spatial position feature and the periodic feature. Wherein, the respective weights of the first spatial position feature and the periodic feature can be determined according to actual test results. For example, the weight corresponding to the first spatial position feature can be θ, and the weight corresponding to the periodic feature can be 1-θ.

[0084] Step 602: determine at least one action feature sample.

[0085] ​The at least one action feature sample is used to represent the respective feature of the at least one action. Specifically, the at least one action feature sample can be determined when the periodic feature extraction network and the spatial feature extraction network are trained. Hereinafter, the model including the periodic feature extraction network, the spatial feature extraction network and the classifier is referred to as an action recognition model. The action recognition model can be trained based on the respective action video of the at least one action, and each action video includes a plurality of image frames corresponding to a complete action.

[0086] In step 603, the action of the target object is recognized based on the action feature of the target object and the at least one action feature sample. The action feature sample is determined based on the plurality of image frames corresponding to the complete action.

[0087] Specifically, the probability that the action of the target object belongs to the action corresponding to the at least one action feature sample can be determined based on the action feature of the target object and the at least one action feature sample, and then the action of the target object is determined based on the probability of each action. For example, the action with the highest probability can be determined as the action of the target object, or the action with a probability greater than a corresponding preset threshold can be determined as the action of the target object.

[0088] It can be understood that for a coherent action such as a traffic police gesture, the historical posture in the historical image frame and the current posture in the target image frame are also important for determining the action of the target object. By using this method, since the action feature sample is determined based on the plurality of image frames corresponding to the complete action, the action feature sample represents the complete action feature, so that the action of the target object can be more accurately determined according to the action feature sample.

[0089] In one possible implementation, as shown in FIG. 7, the following steps 701 to 703 can be performed when the at least one action feature sample is determined. Figure 7

[0090] In step 701, the respective action video of the at least one action is obtained. Each action video includes a plurality of image frames corresponding to a complete action.

[0091] For example, for the stop gesture of a traffic police, the action video corresponding to the gesture should include the complete action from the initial state of the left hand being vertically lowered to the left hand being lifted to the final height.

[0092] In step 702, the first spatial position feature and the periodic feature of the respective action of the at least one action are determined based on the respective action video of the at least one action.

[0093] In step 703, the first spatial position feature and the periodic feature of each action are fused to obtain the action feature sample corresponding to the action.​

[0094] The action recognition model can be trained based on the action videos corresponding to respective actions. For any sample image frame of an action video (i.e., a video frame in the action video), first, the skeletal key points of the target object in the sample image frame are detected based on a human key point detection algorithm. Then, the skeletal key points are input into the action recognition model, and the first spatial position feature corresponding to the sample image frame is extracted by the spatial feature extraction network, and the periodic feature corresponding to the sample image frame is extracted by the periodic feature extraction network. Then, the first spatial position feature and the periodic feature corresponding to the sample image frame are fused according to a preset weight to obtain a fused action feature. Then, the fused action feature is processed by a convolution layer to obtain an action feature sample corresponding to the sample image frame.

[0095] In addition, for any action, after obtaining the fused action feature, the action prediction result can be output by the classifier according to the fused action feature. Then, based on the action prediction result and the label of the action, a loss value can be determined, and then the parameters of the action recognition model are adjusted based on the loss value.

[0096] In a possible implementation, when determining the periodic action corresponding to any sample image frame, the first intermediate feature can be extracted based on the skeletal key points of the target object in the sample image frame; the second intermediate feature can be extracted based on the attention mechanism and the skeletal key points of the target object in the target image frame, and the second intermediate feature includes the second spatial position feature extracted based on the attention mechanism and the skeletal key points of the target object in the historical sample image frame before the sample image frame; and the first intermediate feature and the second intermediate feature are fused to obtain the periodic feature corresponding to the action of the target object.

[0097] It should be noted that the method of determining the periodic feature in the sample image frame can be the same as the method of determining the periodic feature of the target image frame, i.e., extracted based on the first intermediate feature extraction module and the second intermediate feature extraction module, which will not be described here.

[0098] After determining the action feature samples corresponding to respective actions, the action feature samples can be stored in the classifier, so that the classifier can determine the action of the target object in the target image frame based on the action feature samples corresponding to respective actions. By using this method, the action feature samples are determined based on complete action videos, so that the action feature samples have the characteristics of complete actions, and thus the action of the target object in the target image frame can be more accurately determined based on the action feature samples of respective actions.

[0099] It can be understood that different objects of different identities can perform different actions, and the same action can have different meanings for different objects of different identities.

[0100] Therefore, in a possible implementation, after the target image frame is acquired, the identity information of the target object in the target image frame can also be determined first; then when the action of the target object is identified based on the first spatial position feature and the periodic feature, the action of the target object can be determined first based on the first spatial position feature and the periodic feature from a plurality of preset actions corresponding to the identity information.

[0101] The identity represented by the identity information can be a traffic police, a worker, a teacher, or the like. Specifically, preset actions corresponding to each identity information can be set in advance, and then when the action of the target object is determined, the first spatial position feature and the periodic feature are fused to obtain an action feature corresponding to the target object; at least one target action feature sample corresponding to the identity information is determined; and finally, the action of the target object is determined based on the action feature corresponding to the target object and the at least one target action feature sample.

[0102] Alternatively, the first spatial position feature and the periodic feature can be fused first to obtain an action feature corresponding to the target object; at least one action feature sample is determined; the candidate action of the target object is determined based on the action feature corresponding to the target object and the at least one action feature sample; and finally, the action corresponding to the identity information is filtered out from the candidate action as the action of the target object.

[0103] In this way, the action of the target object can be determined according to target objects of different identities, so as to improve the accuracy of the output action.

[0104] In a possible implementation, after the action of the target object in the target image is determined, the prompt information corresponding to the action of the target object can also be displayed.

[0105] Specifically, the prompt information corresponding to at least one action can be set in advance, and after the action of the target object is determined, the prompt information corresponding to the action of the target object can be determined from the prompt information corresponding to at least one action. When the prompt information is displayed, the prompt information can be displayed on a display device, and the prompt information can also be played through an audio playing device. In this way, the user can be prompted in time about the action of the target object.

[0106] In the above embodiments, the target image frame containing the target object can be acquired first, and then based on the target image frame, the skeletal key points of the target object in the target image frame are determined, and then based on the skeletal key points of the target object in the target image frame, the first spatial position feature corresponding to the action of the target object is determined, and the periodic feature corresponding to the action of the target object is determined, and finally the action of the target object is identified based on the first spatial position feature and the periodic feature.

[0107] Since the skeleton key points are not color features of the image itself, they are not easily affected by external environmental factors such as light, and therefore, determining the action feature of the target object based on the skeleton key points can greatly avoid the influence of external environmental factors on the recognition accuracy. In addition, since an action is usually a coherent behavior composed of multiple postures, the way of recognizing the action of the target object in combination with the action features of the target object in the historical image frames can greatly improve the recognition accuracy.

[0108] Based on the same inventive concept, the embodiments of the present application also provide a vehicle control method applied to a terminal (user end) or a server. As shown in Figure 8 Fig. 1 is a flowchart of a vehicle control method provided by an exemplary embodiment of the present application, and specifically, the vehicle control method comprises the following steps 801-803.

[0109] Step 801: determining the action of the target object in the target image based on the action recognition method described in the above embodiments.

[0110] Step 802: determining the vehicle control instruction corresponding to the action of the target object.

[0111] Exemplarily, the vehicle control instruction corresponding to the action of the target object can be determined from at least one vehicle control instruction corresponding to at least one action after the action of the target object is determined.

[0112] Step 803: controlling the target vehicle based on the vehicle control instruction.

[0113] Exemplarily, if the target object is a traffic police, the action of the target object can include a straight gesture, a left turn gesture, a right turn gesture, a stop gesture, etc., and correspondingly, the vehicle control instruction corresponding to the straight gesture can be a straight instruction, and the vehicle control instruction corresponding to the left turn gesture can be a left turn instruction.

[0114] In some implementations, the driving device, the braking device, etc. of the target vehicle are controlled based on the vehicle control instruction to control the driving state (such as the driving speed, the driving direction, etc.) of the target vehicle. For example, when the action of the target object is a left turn gesture, a left turn instruction can be generated, and the driving device of the target vehicle is controlled based on the left turn instruction to make the target vehicle turn left.

[0115] With this method, after the action of the target object is detected, the vehicle can be intelligently controlled according to the action of the target object, without the need for manual control by the driver, which not only saves manpower, but also provides a prerequisite for realizing unmanned driving.

[0116] Exemplary apparatus

[0117] The above describes the method embodiments of the present application in detail. Figures 2 to 8 The device embodiments of the present application are described in detail below in conjunction with Figure 9 and Figure 10 It should be understood that the description of the method embodiments corresponds to the description of the device embodiments, and therefore, the parts not described in detail can be referred to the method embodiments.

[0118] Figure 9 Fig. 1 shows a structural schematic diagram of an action recognition device provided by an example embodiment of the present application. As shown in Fig. 1, the action recognition device 90 provided by the embodiment of the present application comprises: Figure 9

[0119] The obtaining module 901 is configured to obtain a target image frame.

[0120] The first determining module 902 is configured to determine, based on the target image frame, a skeletal key point of a target object in the target image frame.

[0121] The second determining module 903 is configured to determine, based on the skeletal key point of the target object in the target image frame, a first spatial position feature corresponding to an action of the target object.

[0122] The third determining module 904 is configured to determine a periodic feature corresponding to the action of the target object.

[0123] The recognition module 905 is configured to recognize the action of the target object based on the first spatial position feature and the periodic feature.

[0124] In a possible implementation, the third determining module 904 is further configured to: extract a first intermediate feature based on the skeletal key point of the target object in the target image frame; extract a second intermediate feature based on an attention mechanism and the skeletal key point of the target object in the target image frame; and fuse the first intermediate feature and the second intermediate feature to obtain the periodic feature corresponding to the action of the target object.

[0125] In a possible implementation, the recognition module 905 is further configured to: fuse the first spatial position feature and the periodic feature to obtain an action feature corresponding to the target object; determine at least one action feature sample; and recognize the action of the target object based on the action feature corresponding to the target object and the at least one action feature sample.

[0126] In a possible implementation, the recognition module 905 is further configured to: obtain at least one action video corresponding to each of the at least one action, each action video comprising a plurality of image frames corresponding to a complete action; determine, based on the at least one action video corresponding to each of the at least one action, a first spatial position feature and a periodic feature corresponding to each of the at least one action; and fuse the first spatial position feature and the periodic feature corresponding to each of the at least one action to obtain an action feature sample corresponding to the action.​

[0127] In one possible implementation, the first determining module 902 is further configured to: determine the identity information of the target object in the target image frame. Correspondingly, the recognizing module 905 is further configured to: determine the action of the target object from a plurality of preset actions corresponding to the identity information based on the first spatial location features and periodic features.

[0128] Figure 10 The diagram shown is a structural schematic of a vehicle control device provided in an exemplary embodiment of this application. Figure 10 As shown, the vehicle control device 100 provided in this application embodiment includes:

[0129] The first determining module 1001 is used to determine the action of the target object in the target image based on the action recognition method described in the above embodiments;

[0130] The second determining module 1002 is used to determine the vehicle control command corresponding to the action of the target object;

[0131] The control module 1003 is used to control the target vehicle based on vehicle control commands.

[0132] Below, for reference Figure 11 This describes an electronic device according to embodiments of the present application. Figure 11 The diagram shown is a structural schematic of an electronic device provided in an exemplary embodiment of this application.

[0133] like Figure 11 As shown, the electronic device 110 includes one or more processors 1101 and memory 1102.

[0134] The processor 1101 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 110 to perform desired functions.

[0135] The memory 1102 can include one or more computer program products that can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory, for example, can include random access memory (RAM), cache memory, and / or the like. The non-volatile memory, for example, can include read only memory (ROM), hard disk, flash memory, and / or the like. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 1101 can run the program instructions to implement the action recognition method, the vehicle control method, and / or other desired functions of various embodiments of the present application described above. Various contents such as target image frames, skeleton key points, first spatial position features, periodic features, and the like can also be stored in the computer-readable storage media.

[0136] In one example, the electronic device 110 can further include an input device 1103 and an output device 1104, which are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0137] The input device 1103 can include, for example, a keyboard, a mouse, and the like.

[0138] The output device 1104 can output various information including skeleton key points, actions of target objects, and the like to the outside. The output device 1104 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.

[0139] Of course, in order to simplify, Figure 11 Only some of the components of the electronic device 110 related to the present application are shown in the block diagram of FIG. 11, and components such as buses, input / output interfaces, and the like are omitted. In addition, the electronic device 110 can include any other appropriate components according to the specific application.

[0140] In addition to the above-described method and device, an embodiment of the present application can be a computer program product including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the action recognition method, the vehicle control method according to various embodiments of the present application described above in the specification.

[0141] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's computing device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server. The embodiments of the present application are not limited by the programming languages used to write the program code.

[0142] In addition, the embodiments of the present application can also be a computer readable storage medium, which stores computer program instructions, and the computer program instructions make the processor execute the steps of the action recognition method and the vehicle control method according to various embodiments of the present application described above in the specification when the processor runs.

[0143] The computer readable storage medium can take any combination of one or more of the readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can include, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or apparatus, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0144] The above describes the basic principles of the present application in combination with specific embodiments, but it should be noted that the advantages, advantages, effects and the like mentioned in the present application are only examples and are not limited, and these advantages, advantages, effects and the like cannot be considered as the must-have of each embodiment of the present application. In addition, the above specific details are only for the purpose of example and understanding, and the above details do not limit the present application to the must-use specific details.

[0145] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0146] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.

[0147] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0148] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A method of action recognition, characterized by, The method is applied to a vehicle driving scene, and comprises: obtaining a target image frame, the target image frame containing a target object; determining, based on the target image frame, a skeletal key point of the target object in the target image frame; determining, based on the skeletal key point of the target object in the target image frame, a first spatial position feature corresponding to an action of the target object; determining a periodic feature corresponding to the action of the target object, the periodic feature comprising a second spatial position feature corresponding to the action of the target object in a historical image frame before the target image frame; fusing the first spatial position feature and the periodic feature based on respective weights of the first spatial position feature and the periodic feature, to identify the action of the target object; the determining of the periodic feature corresponding to the action of the target object comprises: extracting a first intermediate feature according to a network structure of an LSTM network based on the skeletal key point of the target object in the target image frame, the LSTM network comprising a first intermediate feature extraction module, the first intermediate feature being extracted based on a forgetting gate, an input gate, a time unit, a storage unit and a hidden layer; extracting a second intermediate feature based on an attention mechanism and the skeletal key point of the target object in the target image frame, the second intermediate feature comprising a second spatial position feature extracted based on the attention mechanism and the skeletal key point of the target object in the historical image frame, the LSTM network further comprising a second intermediate feature extraction module, the second intermediate feature being extracted based on a forgetting gate, an input gate, a time-space memory unit and a storage unit, the target image frame being subjected to multi-layer feature processing in the LSTM network to obtain a plurality of levels of second intermediate features; the attention mechanism storing the second spatial position feature corresponding to the historical image frame in the storage unit, the extracting of the second intermediate feature based on the attention mechanism and the skeletal key point of the target object in the target image frame comprising: extracting a third intermediate feature of the skeletal key point of the target object in the target image frame, and then obtaining the second intermediate feature based on the third intermediate feature and the stored second spatial position feature; fusing the first intermediate feature and the second intermediate feature to obtain the periodic feature corresponding to the action of the target object, the second intermediate feature being a last level of second intermediate feature in the LSTM network or the plurality of levels of second intermediate features.

2. The motion recognition method of claim 1, wherein, the identifying of the action of the target object based on the first spatial position feature and the periodic feature comprises: fusing the first spatial position feature and the periodic feature to obtain an action feature corresponding to the target object; determining at least one action feature sample; identifying the action of the target object based on the action feature corresponding to the target object and the at least one action feature sample, the action feature sample being determined based on a plurality of image frames corresponding to a complete action.

3. The motion recognition method of claim 2, wherein, the determining of the at least one action feature sample comprises: acquire at least one action video corresponding to each of the actions, each of the action videos including a plurality of image frames corresponding to a complete action; determine a first spatial position feature and a periodic feature corresponding to each of the at least one action based on the action video corresponding to each of the at least one action; fuse the first spatial position feature and the periodic feature corresponding to each of the actions to obtain an action feature sample corresponding to each of the actions. 4.The motion recognition method of claim 1, wherein, After the target image frame is acquired, the method further includes: determining identity information of the target object in the target image frame; wherein, the action of the target object is identified based on the first spatial position feature and the periodic feature, including: determining the action of the target object from a plurality of preset actions corresponding to the identity information of the target object based on the first spatial position feature and the periodic feature.

5. A vehicle control method characterized by including: determining the action of the target object in the target image based on the action recognition method according to any one of claims 1 to 4; determining a vehicle control instruction corresponding to the action of the target object; controlling a target vehicle based on the vehicle control instruction.

6. An action recognition apparatus characterized by comprising: The device is applied to a vehicle driving scene, including: an acquisition module configured to acquire a target image frame, the target image frame containing a target object; a first determination module configured to determine a skeletal key point of the target object in the target image frame based on the target image frame; a second determination module configured to determine a first spatial position feature corresponding to the action of the target object based on the skeletal key point of the target object in the target image frame; a third determination module configured to determine a periodic feature corresponding to the action of the target object, the periodic feature including a second spatial position feature corresponding to the action of the target object in a historical image frame before the target image frame; an identification module configured to fuse the first spatial position feature and the periodic feature based on respective weights of the first spatial position feature and the periodic feature to identify the action of the target object; the third determination module is further configured to extract a first intermediate feature based on the skeletal key point of the target object in the target image frame according to a network structure of an LSTM network, the LSTM network including a first intermediate feature extraction module, and the first intermediate feature is extracted based on a forgetting gate, an input gate, a time unit, a storage unit and a hidden layer when the first intermediate feature is extracted. The second intermediate feature is extracted based on the attention mechanism and the skeleton key points of the target object in the target image frame, and the second intermediate feature includes a second spatial position feature extracted based on the attention mechanism and the skeleton key points of the target object in the historical image frame, the LSTM network further includes a second intermediate feature extraction module configured to extract the second intermediate feature based on a forgetting gate, an input gate, a space-time memory unit and a storage unit, and the target image frame is subjected to multi-layer feature processing in the LSTM network to obtain a plurality of levels of second intermediate features; the attention mechanism is based on the storage unit storing the second spatial position feature corresponding to the historical image frame, and the second intermediate feature is extracted based on the attention mechanism and the skeleton key points of the target object in the target image frame, including: extracting a third intermediate feature of the skeleton key points of the target object in the target image frame, and then obtaining the second intermediate feature based on the third intermediate feature and the stored second spatial position feature; The first intermediate feature and the second intermediate feature are fused to obtain a periodic feature corresponding to the action of the target object, and the second intermediate feature is the second intermediate feature of the last level in the LSTM network or the second intermediate features of the plurality of levels.

7. A vehicle control device characterized by comprising: Comprising: A first determination module configured to determine an action of a target object in a target image based on the action recognition method of any one of claims 1 to 4; A second determination module configured to determine a vehicle control instruction corresponding to the action of the target object; A control module configured to control a target vehicle based on the vehicle control instruction.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is configured to execute the action recognition method of any one of claims 1 to 4 or the vehicle control method of claim 5.

9. An electronic device, comprising: Comprising: A processor; A memory configured to store instructions executable by the processor; The processor is configured to execute the action recognition method of any one of claims 1 to 4 or the vehicle control method of claim 5.

Citation Information

Patent Citations

  • Gesture recognition method, apparatus, device and operation method based on gesture recognition

    CN104834907A

  • Motion evaluation method and device, equipment and storage medium

    CN112582064A

  • Vehicle control method and device, vehicle and storage medium

    CN114348011A

  • Method, device and equipment for identifying abnormal behaviors in electric power machine room and storage medium

    CN114973097A