An identification method, a training method, a device, a device and a storage medium
By performing three-dimensional and two-dimensional graph convolution processing on the data to be identified in the behavior recognition model, the problem of degradation of recognition accuracy caused by factors such as noise and occlusion is solved, and efficient and accurate pose type recognition is achieved in complex environments.
Patent Information
- Application Number
- CN202011120310.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-19
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-10-19
AI Technical Summary
In actual application, the existing behavior recognition model has greatly reduced the recognition accuracy rate and lacks generalization ability due to factors such as noise, occlusion and data acquisition equipment lag.
A recognition method is adopted to obtain the pose sequence of data to be identified, perform feature extraction, and perform first-level three-dimensional and two-dimensional graph convolution processing to identify the pose type of the target object. The first-level three-dimensional convolution results carry the correlation between local feature information and the first-level two-dimensional convolution results carry the correlation between local feature information, and resist the risks brought by missing information.
It realizes accurate identification of the data to be recognized when the data acquisition device is blocked or stuck, and improves the accuracy and efficiency of posture type recognition in a wider range of application scenarios.
Smart Images

Figure CN114387661B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of information technology, and in particular, to an identification method, a neural network training method, an identification device, a neural network training device, an electronic device, and a computer-readable storage medium. Background Art
[0002] In recent years, behavior recognition technology that recognizes behaviors through collected pose information has been widely applied in fields such as security monitoring, unmanned supermarkets, education and entertainment, intelligent transportation, and smart homes. However, the behavior recognition models used in behavior recognition technology are usually trained with sample data collected under ideal conditions without noise, occlusion, and jitter. In actual behavior recognition scenarios, situations such as noise in data collection, occlusion of pose behaviors, and jitter of collection devices will all strongly interfere with the collected data. At this time, the recognition accuracy of the behavior recognition model trained based on the above sample data for the data to be recognized with a large amount of interference information is greatly reduced. Summary of the Invention
[0003] Embodiments of the present application provide an identification method, a neural network training method, an identification device, a neural network training device, and a computer-readable storage medium.
[0004] The identification method provided by the embodiments of the present application can accurately identify the pose type of the data to be recognized collected under various application scenarios, including when the target object is occluded or the data collection device is jittery, thereby being able to resist the risks brought by the missing information in the data to be recognized, and achieving efficient and accurate identification of the pose type in a wider range of application scenarios.
[0005] The technical solution provided by the embodiments of the present application is as follows:
[0006] Embodiments of the present application provide an identification method, the method including:
[0007] Obtain data to be recognized including at least two pose sequences; wherein, the pose sequence includes multiple node information of any pose of the target object;
[0008] Extract features from the data to be recognized to obtain first-level graph convolution data; wherein, the first-level graph convolution data includes a feature sequence of each pose sequence;
[0009] Perform first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level graph convolution data to obtain a first-level three-dimensional convolution result; wherein, the first-level three-dimensional convolution result includes local feature information in the first-level graph convolution data;
[0010] Perform first-level two-dimensional graph convolution processing on the first-level graph convolution data to obtain a first-level two-dimensional convolution result; wherein, the first-level two-dimensional convolution result includes the association relationship between each of the local feature information in the first-level graph convolution data;
[0011] Based on the first-level graph convolution data, the first-level three-dimensional convolution result, and the first-level two-dimensional convolution result, identify the pose type of the target object.
[0012] An embodiment of the present application further provides a neural network training method. The neural network includes a first network and a second network; wherein, the first network is used for feature extraction; the second network is used for performing graph convolution processing on the output data of the first network; the second network at least includes a first-level three-dimensional graph convolution network for implementing first-level three-dimensional graph convolution processing and a first-level two-dimensional graph convolution network for implementing first-level two-dimensional graph convolution processing; the method includes:
[0013] Obtain sample data including a plurality of pose sequences; wherein, the pose sequence includes a plurality of node information of any pose of any object;
[0014] Based on the first network, perform feature extraction on the sample data to obtain a first-level feature sequence; wherein, the first-level feature sequence includes the feature sequence of each pose sequence;
[0015] Based on the first-level three-dimensional graph convolution network, perform the first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level feature sequence to obtain a first-level three-dimensional result; wherein, the first-level three-dimensional result includes the local feature information in the first-level feature sequence;
[0016] Based on the first-level two-dimensional graph convolution network, perform the first-level two-dimensional graph convolution processing on the first-level feature sequence to obtain a first-level two-dimensional result; wherein, the first-level two-dimensional result includes the association relationship between each of the local feature information in the first feature sequence;
[0017] Based on the first-level feature sequence, the first-level three-dimensional result, and the first-level two-dimensional result, identify the pose type of any object;
[0018] Based on the pose type, train the first network and the second network to obtain a trained neural network.
[0019] An embodiment of the present application further provides an identification device. The identification device includes: a first acquisition module, a first processing module, and a first identification module; wherein:
[0020] The first acquisition module is configured to acquire the data to be recognized including at least two pose sequences; wherein, each pose sequence includes multiple node information of any pose of the target object;
[0021] The first processing module is configured to perform feature extraction on the data to be recognized to obtain first-level graph convolution data; perform first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level graph convolution data to obtain a first-level three-dimensional convolution result; perform first-level two-dimensional graph convolution processing on the first-level graph convolution data to obtain a first-level two-dimensional convolution result; wherein, the first-level graph convolution data includes the feature sequences of each pose sequence; the first-level three-dimensional convolution result includes the local feature information in the first-level graph convolution data; the first-level two-dimensional convolution result includes the correlation between the local feature information in the first-level graph convolution data;
[0022] The first recognition module is configured to recognize the pose type of the target object based on the first-level graph convolution data, the first-level three-dimensional convolution result, and the first-level two-dimensional convolution result.
[0023] An embodiment of the present application further provides a neural network training device. The neural network includes a first network and a second network; wherein, the first network is configured to perform feature extraction; the second network is configured to perform graph convolution processing on the output data of the first network; the second network at least includes a first-level three-dimensional graph convolution network for implementing first-level three-dimensional graph convolution processing and a first-level two-dimensional graph convolution network for implementing first-level two-dimensional graph convolution processing; the device includes a second acquisition module, a second processing module, and a second recognition module; wherein:
[0024] The second acquisition module is configured to acquire sample data including multiple pose sequences; wherein, each pose sequence includes multiple node information of any pose of any object;
[0025] The second processing module is configured to perform feature extraction on the sample data based on the first network to obtain a first-level feature sequence; perform the first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level feature sequence based on the first-level three-dimensional graph convolution network to obtain a first-level three-dimensional result; perform the first-level two-dimensional graph convolution processing on the first-level feature sequence based on the first-level two-dimensional graph convolution network to obtain a first-level two-dimensional result; wherein, the first-level feature sequence includes the feature sequences of each pose sequence; the first-level three-dimensional result includes the local feature information in the first-level feature sequence; the first-level two-dimensional result includes the correlation between the local feature information in the first-level feature sequence;
[0026] The second recognition module is configured to recognize the pose type of any object based on the first-level feature sequence, the first-level three-dimensional result, and the first-level two-dimensional result;
[0027] The second processing module is further configured to train the first network and the second network based on the pose type to obtain a trained neural network.
[0028] An embodiment of the present application further provides an electronic device, which includes a processor, a memory, and a communication bus. The communication bus is used to implement a communication connection between the processor and the memory. The processor is configured to execute a computer program stored in the memory to implement the recognition method as described in any one of the foregoing or the neural network training method as described in any one of the foregoing.
[0029] An embodiment of the present application further provides a computer-readable storage medium that can be executed by a processor to implement the recognition method as described in any one of the foregoing or the neural network training method as described in any one of the foregoing.
[0030] Thus, in the recognition method provided by the embodiment of the present application, first, a feature sequence corresponding to each pose sequence in the data to be recognized is obtained, and then at least two feature sequences in the first-level graph convolutional data are subjected to first-level three-dimensional graph convolution processing to obtain a first-level three-dimensional convolution result including local feature information in the first-level graph convolutional data. The first-level graph convolutional data is subjected to first-level two-dimensional graph convolution processing to obtain a first-level two-dimensional convolution result. Then, based on the first-level graph convolutional data, the first-level three-dimensional convolution result, and the first-level two-dimensional convolution result, the pose type of the target object is recognized.
[0031] As can be seen from the above, in the recognition method provided by the embodiment of the present application, the first-level three-dimensional convolution result carries local feature information in the first-level graph convolutional data, and the first-level two-dimensional convolution result carries the correlation relationship between the local feature information in the first-level graph convolutional data. On this basis, through the first-level three-dimensional convolution result and the first-level two-dimensional convolution result, a set of local feature information with established correlation relationships carried by the data to be recognized can be obtained, that is, the recognition of each dimension of the features carried by the data to be recognized is realized, and a deeper feature recognition process for the data to be recognized is constructed.
[0032] Therefore, even when the data acquisition device is blocked or stuck, that is, the data to be recognized does not carry complete feature data of the target object, the recognition method provided by the present application can still resist the risk brought by missing information based on each local feature information and the correlation relationship between local feature information, and accurately recognize the data to be recognized, thereby realizing efficient and accurate recognition of pose types in a wider range of application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 FIG. 1 is a schematic flowchart of the first recognition method provided by the present application;
[0034] Figure 2a FIG. 2 is a schematic structural diagram of an unoccluded human bone sequence provided by the present application;
[0035] Figure 2b FIG. 3 is a schematic structural diagram of a bone sequence carrying a partial human joint point sequence provided by the present application;
[0036] Figure 2c FIG. 4 is a schematic structural diagram of a bone sequence obtained when a human body translates in a specified direction provided by the present application;
[0037] Figure 3 FIG. 5 is a schematic flowchart of the second recognition method provided by the present application;
[0038] Figure 4 FIG. 6 is a schematic flowchart of feature extraction for each dimension of data in the first data by a multi-layer perceptron provided by the present application;
[0039] Figure 5 FIG. 7 is a schematic flowchart of obtaining the k-th level activation data provided by the present application;
[0040] Figure 6 FIG. 8 is a schematic flowchart of obtaining the k-th level activation data and recognizing the gesture type provided by the present application;
[0041] Figure 7 FIG. 9 is a schematic flowchart of the first neural network training method provided by the present application;
[0042] Figure 8 FIG. 10 is a schematic flowchart of the second neural network training method provided by the present application;
[0043] Figure 9 FIG. 11 is a schematic structural diagram of the recognition device provided by the present application;
[0044] Figure 10 FIG. 12 is a schematic structural diagram of the neural network training device provided by the present application;
[0045] Figure 11 FIG. 13 is a schematic structural diagram of the behavior recognition system provided by the present application;
[0046] Figure 12 FIG. 14 is a schematic structural diagram of the electronic device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application.
[0048] It should be understood that the specific embodiments described herein are merely for explaining the present application and are not used to limit the present application.
[0049] The present application relates to the field of information technology, and in particular, to an identification method, a neural network training method, an identification device, a neural network training device, and a computer-readable storage medium.
[0050] In recent years, behavior recognition technology has been widely applied in fields such as security monitoring, unmanned supermarkets, education and entertainment, intelligent transportation, and smart homes. The application of behavior recognition technology has also greatly improved the intelligence level of monitoring devices.
[0051] The main research objective of behavior recognition technology is to analyze the action behaviors of objects in a visual scene through data collected by various types of sensors.
[0052] In the related art, behavior recognition technology usually includes the following methods:
[0053] Directly perform behavior recognition based on the data collected by sensors, the images or video data directly collected by cameras; extract skeletal data from the data collected by infrared cameras and millimeter-wave radars, and then perform behavior recognition based on the skeletal data.
[0054] In some Internet of Things scenarios with high privacy requirements, high communication delays, and limited computing resources, the behavior recognition method based on skeletal data has obvious advantages. For example, in the human-computer interaction scenario of smart homes, users' acceptance of extracting human skeletons from the data collected by infrared depth cameras and then performing behavior recognition based on the human skeletons is higher than that of recognizing the live videos collected by ordinary cameras; in the monitoring scenario of intelligent transportation, after extracting skeletal data from the image or video data collected by cameras and then transmitting the skeletal data to the cloud for behavior recognition, it can also greatly reduce the network data transmission pressure, and thus can reduce the data transmission delay.
[0055] However, the behavior recognition model sampled in the behavior recognition technology in the related art is trained through sample data without any noise and interference. Such sample data that meets the experimental requirements is very ideal for the training of the behavior recognition model. However, the behavior recognition model trained based on this sample data has a high theoretical value of behavior recognition during algorithm simulation, but the behavior recognition effect in actual behavior recognition applications is poor. For example, for the recognition of jumping actions, the behavior recognition model trained based on the above sample data can achieve a test accuracy rate of over 90% for the standard test data set. However, in actual applications, due to factors such as camera jitter, obstacle occlusion, and shooting angle, the recognition accuracy rate of this behavior recognition model is even lower than 60%.
[0056] This is because the target data corresponding to the above behavior recognition method can only be sample data without noise and interference that is the same as the sample data, that is, it can overfit clean data, and it lacks the ability to generalize data with various noises and interferences in actual applications, thus greatly reducing its data recognition performance in actual applications.
[0057] Based on the above problems, the embodiment of the present application provides a recognition method, which can be implemented by a processor of a recognition device. The above processor can be at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor.
[0058] Figure 1 It is a schematic flowchart of the first recognition method provided by the embodiment of the present application. The recognition method may include the following steps:
[0059] Step 101, obtain the data to be recognized including at least two pose sequences.
[0060] Wherein, the pose sequence includes multiple node information of any pose of the target object.
[0061] In one embodiment, the target object can be an object without vital signs, such as a tree swayed by the wind, a falling leaf, a flowing river, a moving sand dune, etc.
[0062] In one embodiment, the target object can be an object without vital signs but capable of changing its motion state by receiving control instructions or external driving forces, such as a robot, an electronic toy, an electric vehicle, etc.
[0063] In one embodiment, the target object can be an object with vital signs, such as a person, any type of animal, etc.
[0064] In one embodiment, the target object can be an object that appears in a specific scenario, such as a stranger who appears in a sensitive area of a bank.
[0065] In one embodiment, the target object can include multiple different types of objects. For example, a pet dog being led and a pedestrian walking the pet dog.
[0066] In one embodiment, the target object can be a part of any object, such as the crown of a big tree, the top of a sand dune, the wheel of an electric vehicle, the arm of a pedestrian, the limbs of a pet, etc.
[0067] In one embodiment, the target object can include different parts of multiple objects, such as the left foot of the first person and the right foot of the second person.
[0068] In one embodiment, the target object can be partially blocked by an obstacle.
[0069] In one embodiment, any posture of the target object can be any posture of the target object in a static state or in a motion state.
[0070] In one embodiment, any posture of the target object can represent any posture of the target object during the process of switching from a static state to a motion state or from a motion state to a static state.
[0071] In one embodiment, the posture sequence can represent a data sequence obtained by quantifying any posture of the target object in a specified coordinate system. For example, when the target object is in a static state, the posture sequence can represent the coordinate sequence of each joint of the target object's leg in the specified coordinate system.
[0072] In one embodiment, the posture sequence can represent all or part of the bone sequence corresponding to any posture of the target object. Exemplarily, the partial bone sequence can represent the bone sequence that is not blocked by an obstacle.
[0073] In one implementation, a bone sequence may include a sequence of coordinates of key points in each part of a target object. Exemplarily, the bone sequence may represent a sequence of each key joint of a human body.
[0074] Exemplarily, Figure 2a FIG. is a schematic structural diagram of an unobstructed human bone sequence provided by an embodiment of the present application.
[0075] In Figure 2a the human bone sequence includes sequence points corresponding to various parts such as the limbs, torso, head, and cervical vertebrae.
[0076] In Figure 2a the bone sequence may include sequence points of each joint of the human body.
[0077] In one implementation, the bone sequence may include sequence points of some joints of the human body. Exemplarily, Figure 2b FIG. is a schematic structural diagram of a bone sequence carrying a partial human joint point sequence provided by an embodiment of the present application. In Figure 2b it, the gray part represents the joint points of the human body part blocked by an obstacle, and the white part is the joint points of the human body part not blocked by the obstacle.
[0078] Exemplarily, when the human body is in a moving state, the state of the bone sequence will also change. Figure 2c FIG. is a schematic structural diagram of a bone sequence obtained when the human body translates in a specified direction provided by an embodiment of the present application. In Figure 2c it, the gray part represents the joint point sequence blocked when the human body translates, and the white part represents the joint point sequence not blocked when the human body translates.
[0079] Exemplarily, Figure 2c the bone sequence shown in may represent a bone sequence obtained from data collected when the data acquisition device freezes. Among them, the data acquisition device may include at least one of the following: various types of image acquisition devices, sensor devices, ranging devices, etc.
[0080] In one implementation, the data to be recognized may be any number of pose sequences in a pose sequence set. Exemplarily, the pose sequence set may include a large number of pose sequences.
[0081] In one implementation, the data to be recognized may be multiple temporally continuous pose sequences obtained from the pose sequence set in chronological order after sorting each pose sequence in the pose sequence set according to time.
[0082] In one embodiment, the data to be recognized may include multiple pose sequences corresponding to the acquisition results of the same pose information of a target object by multiple data acquisition devices at different angles or positions.
[0083] In one embodiment, the data to be recognized may include the pose sequence of one target object or the pose sequences of multiple target objects.
[0084] Step 102: Extract features from the data to be recognized to obtain first-level graph convolution data.
[0085] Among them, the first-level graph convolution data includes the feature sequence of each pose sequence.
[0086] In one embodiment, extracting features from the data to be recognized may be to remove the noise and interference in the data to be recognized to obtain the effective feature data carried in the data to be recognized.
[0087] Correspondingly, the first-level graph convolution data may include the effective feature sequence in each pose sequence.
[0088] In one embodiment, extracting features from the data to be recognized may be to extract partial node information in the data to be recognized.
[0089] Correspondingly, the first-level graph convolution data may include the feature sequence corresponding to partial nodes in each pose sequence.
[0090] In one embodiment, extracting features from the data to be recognized may be to perform linear feature extraction on any pose sequence in the data to be recognized.
[0091] In one real-time mode, extracting features from the data to be recognized may be performed in a non-linear manner.
[0092] In one embodiment, extracting features from the data to be recognized may be implemented by a perceptron. Exemplarily, the perceptron may be multi-layer.
[0093] In one embodiment, the first-level graph convolution data may include at least two feature sequences.
[0094] Step 103: Perform first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level graph convolution data to obtain a first-level three-dimensional convolution result.
[0095] Among them, the first-level three-dimensional convolution result includes the local feature information in the first-level graph convolution data.
[0096] In one embodiment, the local feature information in the first-level graph convolution data may represent the feature information of a specified part of each feature sequence in the first-level graph convolution data.
[0097] In one embodiment, the local feature information in the first-level graph convolution data may represent the feature information of a specified feature sequence in the first-level graph convolution data.
[0098] In one embodiment, when the data to be recognized and the target object are blocked by an obstacle, or the data collected by the data acquisition device is stuck, the local feature information in the first-level graph convolution data may represent all the feature information carried by each feature sequence in the first-level graph convolution data.
[0099] In one embodiment, the first-level three-dimensional graph convolution process may be to perform graph convolution processing on at least two feature sequences in the first-level graph convolution data in sequence, and perform splicing analysis on the results of the graph convolution processing to obtain the first-level three-dimensional convolution result.
[0100] In one embodiment, the first-level three-dimensional graph convolution process may be to analyze, splice, and summarize each data point in at least two feature sequences in the first-level graph convolution data to obtain an intermediate result, and perform graph convolution processing on the intermediate result to obtain the first-level three-dimensional convolution result.
[0101] In one embodiment, the first-level three-dimensional graph convolution process may be to analyze each data point in at least two feature sequences in the first-level graph convolution data to obtain a first node set including each of the above data points, a first edge set corresponding to the first node set, and then obtain a first adjacency matrix corresponding to the first edge set, and calculate the three-dimensional graph convolution based on the first adjacency matrix.
[0102] Exemplarily, since each data point in the first node set combination includes each data point in at least two feature sequences, and at least two feature sequences may include feature sequences collected by the target object at different times or by multiple data acquisition devices at multiple acquisition angles, each data point included in the first node set can represent the change of the pose information of the target object over time or over space. That is to say, the first node set can represent the change information of the pose of the target object in three-dimensional space.
[0103] Step 104: Perform first-level two-dimensional graph convolution processing on the first-level graph convolution data to obtain a first-level two-dimensional convolution result.
[0104] Among them, the first-level two-dimensional convolution result includes the correlation relationship between the local feature information in the first-level graph convolution data.
[0105] In one embodiment, the association relationship between local feature information in the first-level graph convolution data can indicate whether there is an association relationship between local feature information in the first-level graph convolution data.
[0106] In one embodiment, the association relationship between local feature information in the first-level graph convolution data can indicate the strength of the association relationship between local feature information in the first-level graph convolution data.
[0107] In one embodiment, the association relationship between local feature information in the first-level graph convolution data can indicate the association relationship between local feature information carried in each individual sequence in the first-level graph convolution data.
[0108] In one embodiment, the first-level two-dimensional graph convolution process can be to analyze any sequence in the first-level graph convolution data to obtain a second node set and a second edge set, and then obtain a second adjacency matrix corresponding to the second node set, and perform graph convolution calculation based on the second adjacency matrix.
[0109] In one embodiment, the first-level two-dimensional graph convolution process can be an operation performed sequentially on each sequence in the first-level graph convolution data.
[0110] It should be noted that under the condition that the recognition device has a multi-core processor, step 103 and step 104 can be performed in parallel, or the execution order of step 103 and step 104 can be adjusted, and the embodiments of the present application do not limit this.
[0111] Step 105, based on the first-level graph convolution data, the first-level three-dimensional convolution result, and the first-level two-dimensional convolution result, identify the pose type of the target object.
[0112] In one embodiment, the pose type of the target object can represent the state type corresponding to the current pose of the target object. Exemplarily, the state type can include a stationary state and a motion state.
[0113] In one embodiment, the pose type of the target object can represent the motion state type corresponding to the current pose of the target object. Exemplarily, the motion state type can include at least one of accelerated motion, decelerated motion, uniform motion, switching from a stationary state to a motion state type, switching from a motion state to a stationary state type, etc.
[0114] In one implementation, the posture type of the target object can indicate whether the current posture of the target object matches a preset posture range. Exemplarily, the posture range can correspond to a posture recognition scenario. For example, it includes the interaction modes between intelligent electronic devices and users, such as unmanned supermarkets, education and entertainment, and smart homes. The interaction mode can include at least one of the following: gesture interaction, body interaction, relative motion direction interaction, and collective interaction of the above several methods.
[0115] In one implementation, the posture type of the target object can represent the relative position change or relative speed change between the current posture of the target object and other objects. Exemplarily, in the scenario of security monitoring, by detecting the relative position change or relative speed change between the target object and other objects, it can be used as a basis for analyzing the legal facts corresponding to a certain posture of the target object.
[0116] In one implementation, the posture type of the target object can also be used in competitive sports to assist in judging whether the posture of the target object is standard.
[0117] In one implementation, the posture type of the target object can be obtained by the following method:
[0118] Identify the posture type of the target object according to the intersection between the first-level graph convolution data and the results of the first-level three-dimensional convolution and the first-level two-dimensional convolution.
[0119] In one implementation, the posture type of the target object can be obtained by the following method:
[0120] Combine the local feature information in the first-level graph convolution data carried by the first-level three-dimensional convolution result and the correlation relationship between the local feature information carried by the first-level two-dimensional convolution results, associate the local feature information to obtain an association result, then match the association result with the first-level graph convolution data to obtain a matching result, and identify the posture type of the target object according to the matching result.
[0121] The association result obtained through the above operations can resist the interference to the data to be recognized caused by various factors in actual applications, such as the data acquisition device being blocked, the data acquisition device being stuck, the installation angle of the data acquisition device changing, and the target object being blocked. Furthermore, the recognition method provided by the embodiments of the present application can stably and efficiently recognize various data to be recognized in actual applications.
[0122] Thus, in the recognition method provided by the embodiments of the present application, first, the feature sequence corresponding to each pose sequence in the data to be recognized is obtained, and then at least two feature sequences in the first-level graph convolution data are subjected to first-level three-dimensional graph convolution processing to obtain a first-level three-dimensional convolution result containing local feature information in the first-level graph convolution data. The first-level graph convolution data is subjected to first-level two-dimensional graph convolution processing to obtain a first-level two-dimensional convolution result containing the correlation relationship between local feature information in the first graph convolution data. Then, based on the first-level graph convolution data, the first-level three-dimensional convolution result, and the first-level two-dimensional convolution result, the pose type of the target object is recognized.
[0123] As can be seen from the above, in the recognition method provided by the embodiments of the present application, the local feature information in the first-level graph convolution data is carried in the first-level three-dimensional graph convolution result, and the correlation relationship between local feature information in the first-level graph convolution data is carried in the first-level two-dimensional graph convolution result. On this basis, by combining the first-level three-dimensional convolution result and the first-level two-dimensional convolution result, a set of local feature information with established correlation relationships carried by the data to be recognized can be obtained, that is, the recognition of each dimension of the features carried by the data to be recognized is realized, and a deeper feature recognition process for the data to be recognized is constructed. Thus, even when the data acquisition device is blocked or stuck, that is, the data to be recognized does not carry complete feature data of the target object, the recognition method provided by the present application can still resist the risk brought by missing information based on each local feature information and the correlation relationship between local feature information, and accurately recognize the data to be recognized, thereby realizing the efficient and accurate recognition of the pose type in a wider application scenario.
[0124] Based on the foregoing embodiments, the embodiments of the present application provide a second recognition method. Figure 3 It is a schematic flowchart of the second recognition method provided by the embodiments of the present application. As Figure 3 shown, the recognition method may include the following steps:
[0125] Step 301, obtain the data to be recognized including at least two pose sequences.
[0126] Wherein, the pose sequence includes multiple node information of any pose of the target object.
[0127] Step 302, perform feature extraction on the data to be recognized to obtain first-level graph convolution data.
[0128] Wherein, the first-level graph convolution data includes the feature sequence of each pose sequence.
[0129] Exemplarily, step 302 can be implemented through steps A1 - A3:
[0130] Step A1: Obtain the first data from the data to be recognized.
[0131] Among them, the first data represents data of at least two dimensions corresponding to the data to be recognized.
[0132] In one implementation, the first data can be obtained by performing multi-dimensional decomposition on each data point in the data to be recognized.
[0133] In one implementation, the first data can be obtained by mapping each data point in the data to be recognized to a coordinate system of a specified dimension and based on the values of each dimension of each data point in the coordinate system of the specified dimension.
[0134] In one implementation, under the condition that the data to be recognized represents a bone sequence, the data to be recognized can be represented by Equation (1):
[0135] X = {p t,i |t = 1,…T; i = 1,…V} (1)
[0136] In Equation (1), X represents the bone sequence, t represents the number of frames corresponding to the bone sequence. For multiple bone sequences, t can also be used to represent the time corresponding to the acquisition of the bone sequence, and its maximum value is T; i represents the number of nodes included in each bone sequence, and its maximum value is V; p t,i represents the i-th node in the t-th frame of the bone sequence. Among them, both V and T are integers greater than or equal to 2.
[0137] Under the condition of mapping the bone sequence in a three-dimensional coordinate system, the position of the bone node p t,i can be represented by Equation (2):
[0138] p t,i = (x t,i , y t,i , z t,i ) T (2)
[0139] Among them, x t,i , y t,i , z t,i are respectively used to represent the coordinate values corresponding to the three coordinate axes of p t,i in the three-dimensional space.
[0140] Exemplarily, the above at least two dimensions may include at least two of the following dimensions: distance dimension, speed dimension, and position dimension.
[0141] In one embodiment, the distance dimension may represent the distance of a certain bone node relative to other bone nodes. Under the condition that there is a connection relationship between two bone nodes, the distance dimension may represent the length of the backbone connecting the two bone nodes.
[0142] In one embodiment, the distance dimension may represent the distance of a certain bone node relative to a specified bone node.
[0143] Exemplarily, bone node p t,i and bone node p t,i-1 The value b of the distance dimension between t,i , can be calculated by Equation (3).
[0144] b t,i = p t,i - p t,i-1 (3)
[0145] In Equation (3), there is a connection relationship between bone node p t,i and bone node p t,i-1 . Therefore, b t,i can be the backbone connecting the above two bone nodes. Wherein, i is an integer greater than 1.
[0146] In one embodiment, the speed dimension may represent the amount of change in distance per unit time of a certain bone node relative to the origin of the coordinate system.
[0147] Exemplarily, the value v of the speed dimension of bone node p t,i can be calculated by Equation (4). t,i
[0148] v t,i = p t,i - p t-1,i (4)
[0149] Exemplarily, v t,i can represent the speed of bone node p t,i .
[0150] Step A2: Perform feature extraction on the data of each dimension in the first data to obtain the second data.
[0151] In one embodiment, performing feature extraction on the data of each dimension in the first data may be implemented by a perceptron. Among them, the perceptron can be multi-layer.
[0152] Figure 4 This is a schematic flow diagram of performing feature extraction on the data of each dimension in the first data by a multi-layer perceptron provided by the embodiments of the present application.
[0153] As Figure 4 shown, after obtaining the data of the position dimension, distance dimension, and speed dimension respectively from the data to be recognized containing two skeletal sequences, the data of the position dimension, distance dimension, and speed dimension are respectively input into the first multi-layer perceptron, the second multi-layer perceptron, and the third multi-layer perceptron. Among them, the first multi-layer perceptron, the second multi-layer perceptron, and the third multi-layer perceptron can extract features from the input data of each dimension through a non-linear function.
[0154] Exemplarily, taking the distance dimension b t,i in the above text as an example, under the condition that the second multi-layer perceptron has two layers, the feature extraction process of the second multi-layer perceptron can be realized by formula (5):
[0155]
[0156] In formula (5), W1 and W2 are respectively used to represent the weights of the first feature extraction layer and the second feature extraction layer of the second multi-layer perceptron, and σ() is used to represent the activation function, which is usually a non-linear function; is used to represent the result of the second multi-layer perceptron's feature extraction of b t,i feature.
[0157] Since the process of extracting features from the data of each dimension through the above multiple multi-layer perceptrons is equivalent to further extracting and quantifying the data features of each dimension, therefore, the feature extraction process of the above each multi-layer perceptron can also be called the process of feature encoding the data of each dimension.
[0158] Exemplarily, the results of the above multiple multi-layer perceptrons extracting features from the data of each dimension can be respectively recorded as position features, distance features, and speed features.
[0159] In one implementation, the second data can include the results of feature extraction for the data of each dimension respectively. As Figure 4 shown, the second data can include position features, distance features, and speed features.
[0160] The feature extraction process for the data of multiple dimensions can make full use of the effective information carried by the features of each dimension, thereby laying a foundation for obtaining richer dependency relationships and internal connections in the subsequent posture type recognition process, and is conducive to weakening the data drift caused by noise interference in the data to be recognized. In this way, when a large amount of noise or interference exists in any dimension of the data to be recognized, the feature information of the data to be recognized can still be obtained through other dimensions. That is to say, through the feature extraction of different dimensions of the data to be recognized, the feature extraction results of each dimension can complement and calibrate each other, and can absorb and weaken the influence of perturbations and interferences.
[0161] Step A3: Obtain the first-level graph convolutional data based on the second data.
[0162] Exemplarily, the first-level graph convolutional data can be obtained by fusing the feature data of each dimension in the second data.
[0163] In one implementation, the first-level graph convolutional data can be obtained by processing the second data through feature aggregation.
[0164] Exemplarily, feature aggregation can be achieved through a trained feature training model.
[0165] In one implementation, the first-level graph convolutional data can be obtained by using the mean method to fuse each feature data in the second data. Exemplarily, the mean method can be to perform an average combination of the feature information of three dimensions.
[0166] In one implementation, the first-level graph convolutional data can be obtained by fusing each feature data in the second data through the concatenation method. Exemplarily, the concatenation method can be to concatenate the feature information of the above three dimensions together end to end.
[0167] Through the above operations, after aggregating the feature information of each extracted dimension, the first-level graph convolutional data obtained carries more comprehensive feature information.
[0168] Step 303: Perform first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level graph convolutional data to obtain a first-level three-dimensional convolution result.
[0169] Among them, the first-level three-dimensional convolution result includes the local feature information in the first-level graph convolutional data.
[0170] Exemplarily, Step 303 can be implemented through Steps B1 - B4:
[0171] Step B1: Determine the sliding window.
[0172] In one implementation, the sliding window can represent the length information corresponding to the time dimension. Correspondingly, the sliding window can represent a 5-second time window.
[0173] In one implementation, the sliding window can represent the quantity information of data selection. Correspondingly, the sliding window can represent selecting 5 data at a time.
[0174] In one implementation, the sliding window can be selected according to the refresh frequency of the data to be selected.
[0175] In one embodiment, the sliding window can be set according to the accuracy of data analysis.
[0176] In one embodiment, the sliding window can be adjusted according to different data analysis methods.
[0177] Step B2: Based on the sliding window, obtain at least two feature sequences from the first-level graph convolution data.
[0178] In one embodiment, the at least two feature sequences can be obtained in the following way:
[0179] Based on the sliding window, arrange each feature sequence in the first-level graph convolution data in a specified order, and select at least two feature sequences from the arrangement result. Exemplarily, the specified order can represent the time order.
[0180] Exemplarily, under the condition that the sliding window represents the number of data selected from the first graph convolution data, if the value of the sliding window is τ, the number of feature sequences that can be obtained from the first graph convolution data through one sliding window can be τ 。 where τ is an integer greater than 1.
[0181] Correspondingly, when the number of skeleton nodes included in each feature sequence is N, the number of skeleton nodes obtained through one sliding window τ is Nτ 。
[0182] Step B3: Determine the first-level three-dimensional data based on the at least two feature sequences.
[0183] Among them, the first-level three-dimensional data includes each node in the at least two feature sequences and the association relationship between any two nodes in the at least two feature sequences.
[0184] In one embodiment, each node in the at least two feature sequences can be embodied in the form of a first node set.
[0185] In one embodiment, for the convenience of counting each node in the at least two feature sequences, the first node set can be summarized in the form of a matrix. Exemplarily, the first node set can be denoted as V τ .
[0186] In one embodiment, the association relationship between any two nodes in the at least two feature sequences can represent whether there is an association relationship between any two nodes.
[0187] In one embodiment, the association relationship between any two nodes in the at least two feature sequences can represent the strength of the association between any two nodes.
[0188] Exemplarily, a bone sequence can be represented by an undirected graph. Correspondingly, the association relationship between any two bone nodes can be reflected by whether there is an association relationship between any two bone nodes.
[0189] Exemplarily, when the number of bone nodes is Nτ, the association relationship between any two bone nodes can be represented by an Nτ×Nτ matrix, and this matrix can be the first adjacency matrix A τ ,n, and can be represented by Equation (6):
[0190]
[0191] In Equation (6), a 0,0 is used to represent the association relationship between the 0th bone node of the first bone sequence and the 0th bone node of the second bone sequence; a 0,τ-1 is used to represent the association relationship between the 0th bone node of the first bone sequence and the (τ - 1)th node of the second bone sequence; a N-1,0 is used to represent the association relationship between the (N - 1)th bone node of the first bone sequence and the 0th bone node of the second bone sequence; a N-1,τ-1 is used to represent the association relationship between the (N - 1)th bone node of the first bone sequence and the (τ - 1)th bone node of the second bone sequence.
[0192] The first adjacency matrix A τ ,n shown in Equation (6) 0,0 The value of each element such as a 0,τ-1 , a 0,τ-1 can be 0 or 1. Exemplarily, if any element value is 0, it means there is no association relationship between the two corresponding bone nodes, that is, there is no connection between the two bone nodes; if any element value is 1, it means there is an association relationship between the two corresponding bone nodes, that is, the two bone nodes are connected to each other.
[0193] In one implementation, the first - level three - dimensional data can include the first node set V τ and the first adjacency matrix A τ,n .
[0194] In one implementation, the first - level three - dimensional data can be obtained by splicing the first node set V τ and the first edge set E τ in a certain method.
[0195] Exemplarily, the first - level three - dimensional data can be denoted as G τ , and the first - level three - dimensional data can be represented by Equation (7):
[0196] G τ = (V τ , E τ ) (7)
[0197] Step B4: Perform first-level 3D graph convolution processing on the first-level 3D data.
[0198] Exemplarily, performing first-level 3D graph convolution processing on the first-level 3D data can be achieved through Equation (8):
[0199]
[0200] where n ∈ [0, N] is the neighbor node range, A τ,n is the first adjacency matrix of size Nτ × Nτ; is the degree matrix corresponding to the first adjacency matrix A τ,n ; is the weight matrix of the first-level 3D graph convolution processing; is the first-level 3D data; is the semantic matrix representing the implicit relationship between the adjustable skeleton nodes. Exemplarily, it can be a parameter that can be updated iteratively by backpropagation according to the neural network recognition loss; is the output data of the first-level 3D graph convolution processing, that is, the first-level 3D convolution result. Here, N is an integer greater than 1.
[0201] Through the first-level 3D graph convolution processing, local feature information can be fully extracted. Through these local feature information, motion details can be focused on, which is beneficial to accurately judging the motion category.
[0202] It should be noted that the recognition method provided in the embodiments of the present application can be implemented by a neural network completed through training. The above various parameters can represent the parameters of the neural network.
[0203] Step 304: Perform first-level 2D graph convolution processing on the first-level graph convolution data to obtain the first-level 2D convolution result.
[0204] Among them, the first-level 2D convolution result includes the association relationship between the local feature information in the first-level graph convolution data.
[0205] It should be noted that the operation of the first-level 2D graph convolution processing is similar to that of the first-level 3D graph convolution, both of which are performed through Equation (8).
[0206] Different from the process of the first-level 3D graph convolution processing, the input data of the first-level 2D graph convolution processing is the topological graph of any frame of skeleton nodes in the first-level graph convolution data, rather than at least two frames of data.
[0207] Correspondingly, for the first-level two-dimensional graph convolution processing, a topological graph can be first established for a single-frame bone sequence, and two-dimensional graph convolution operations can be performed to extract the spatial features between each node.
[0208] Exemplarily, after the first-level two-dimensional graph convolution processing, a first-level lightweight convolution operation can also be performed on the result of the first-level two-dimensional graph convolution processing.
[0209] Exemplarily, the first-level lightweight convolution operation can decompose an ordinary convolution into convolutions in multiple directions, and superimpose the convolution results in multiple directions. Figure 1 The first lightweight convolution operation shown can be decomposed into a convolution in the depth direction and a convolution in the width direction.
[0210] Through the k-th level lightweight convolution operation, the number of parameters that need to be adjusted during the training of the neural network can be reduced, enabling the neural network to be lightweight, and also enabling as much effective information as possible to be extracted with a relatively small computational cost, thereby facilitating the deployment of the neural network to improve the computational efficiency of data.
[0211] Moreover, through the first-level two-dimensional graph convolution processing and the first-level lightweight convolution operation, general features in the data to be recognized can be extracted, which is also beneficial for the deployment of recognition devices in scenarios with limited resources in the Internet of Things.
[0212] Step 305: Fuse the first-level three-dimensional convolution result and the first-level two-dimensional convolution result to obtain the first-level activation data.
[0213] Among them, the first-level activation data represents the feature information recognized from the first-level graph convolution data.
[0214] In one implementation, the first-level activation data can be obtained by performing feature splicing on the first-level three-dimensional convolution result and the first-level two-dimensional convolution result.
[0215] In one implementation, the first-level activation data can be realized by performing lightweight convolutions in multiple dimensions on the first-level three-dimensional convolution result and the first-level two-dimensional convolution result.
[0216] Figure 5 The figure shows a schematic flowchart of obtaining the k-th level activation data provided by the embodiments of the present application.
[0217] Exemplarily, in Figure 5Among them, k is an integer greater than or equal to 1. Under the condition that k is 1, for the first-level activation data, at least two feature sequences are obtained from the first-level graph convolution data through a sliding window, and through the process of the foregoing embodiments, at least two feature sequences are subjected to first-level three-dimensional graph convolution processing to obtain a first-level three-dimensional convolution result; at the same time, the first-level graph convolution data is subjected to first-level two-dimensional graph convolution processing to obtain a first-level two-dimensional convolution result; then the first-level two-dimensional convolution result and the first-level three-dimensional convolution result are subjected to fusion processing to obtain the first-level activation data.
[0218] Exemplarily, the first-level activation data can be obtained through Equation (9).
[0219]
[0220] In Equation (9), f n (x, y) represents the result of fusing the first-level three-dimensional convolution data and the first-level two-dimensional convolution result, denoted as the first-level graph convolution data fusion result; ω n,c is the k-channel weight corresponding to the c-th type of pose; M c (x, y) is the first activation data of the c-th type of pose, which can be abbreviated as M c,s . Among them, c is an integer greater than 1.
[0221] Figure 6 This is a schematic flowchart of obtaining the k-th level activation data and identifying the pose type provided by the embodiments of the present application.
[0222] Under the condition that k is 1, Figure 6 As shown, it is a flowchart of obtaining the first-level activation data and identifying the pose type.
[0223] In Figure 6 , based on the first-level graph convolution data fusion result, a visual first-level feature activation map can be obtained, that is, the first-level activation data is obtained, and the pose type is identified based on this activation data.
[0224] Exemplarily, Figure 6 as shown, after obtaining the first-level activation data based on the first-level graph convolution data fusion result and the first-level graph convolution data, the first-level under-activated or first-level unactivated data in the first-level graph convolution data can also be obtained. Specifically, these first-level under-activated or first-level unactivated data can be realized through Equation (10):
[0225] mask s = 1 - softmax(M c,s ) (10)
[0226] In Equation (10), mask sUsed to represent the first-level under-activated or first-level unactivated data obtained after weighted calculation of the first activation data M by the softmax() function. c,s After performing weighted calculation on
[0227] Step 306: Identify the pose type based on the first-level activation data and the first-level graph convolution data.
[0228] In one implementation, the pose type can be determined based on the gap between the first-level activation data and the first-level graph convolution data, and the pose type can be identified according to this method.
[0229] Exemplarily, step 306 can be implemented through steps C1 - C3:
[0230] Step C1: Determine the k-th level graph convolution data based on the (k - 1)-th level activation data and the (k - 1)-th level graph convolution data.
[0231] Where k is an integer greater than 1; the k-th level graph convolution data includes the feature sequences not recognized in the (k - 1)-th level graph convolution data.
[0232] Exemplarily, for the feature data not recognized in the first-level two-dimensional graph convolution processing and the first-level three-dimensional graph convolution processing, it can be input into the next-level two-dimensional graph convolution processing and the next-level three-dimensional graph convolution processing for processing.
[0233] Exemplarily, the k-th level graph convolution data can be obtained through Equation (11):
[0234]
[0235] In Equation (11), is used to represent the (k - 1)-th level graph convolution data; is used to represent the k-th level graph convolution data. When k is 0, is used to represent the first-level graph convolution data in the foregoing embodiments.
[0236] Step C2: Obtain the k-th level activation data based on the k-th level graph convolution data.
[0237] Where the k-th level activation data represents the feature information recognized from the k-th level graph convolution data.
[0238] Exemplarily, the k-th level activation data can be implemented through Equation (9).
[0239] Exemplarily, step C2 can be implemented through steps D1 - D3:
[0240] Step D1: Perform k-th level 3D graph convolution processing on at least two feature sequences in the k-th level graph convolution data to obtain the k-th level 3D convolution result.
[0241] Exemplarily, the k-th level 3D graph convolution processing can be implemented by Equation (12):
[0242]
[0243] In Equation (12), is used to represent the output data of the k-th level 3D graph convolution processing, i.e., the k-th level 3D convolution result; is the weight matrix corresponding to the k-th level 3D graph convolution processing; is the concatenation result of the input data of the k-th level 3D graph convolution processing, i.e., the matrix corresponding to the k-th node set and the matrix corresponding to the k-th edge set of the k-th level graph convolution data.
[0244] Exemplarily, it can be to first determine the k-th level 3D data, and then perform the k-th level 3D graph convolution processing based on the k-th level 3D data to obtain the k-th level 3D convolution result.
[0245] Step D2: Perform k-th level 2D graph convolution processing on the k-th level graph convolution data to obtain the k-th level 2D convolution result.
[0246] The process of the k-th level 2D graph convolution processing is the same as that of the first level 2D graph convolution processing, which will not be elaborated here.
[0247] Step D3: Based on the k-th level 3D convolution result and the k-th level 2D convolution result, obtain the k-th level activation data.
[0248] The process of obtaining the k-th level activation data is the same as that of obtaining the first level activation data, which will not be elaborated here.
[0249] Step C3: Based on the first level activation data to the k-th level activation data, identify the pose type.
[0250] Exemplarily, the recognition method provided in the embodiments of the present application can be implemented through 3D graph convolution processing at multiple levels and 2D graph convolution processing at multiple levels. Correspondingly, k 3D convolution results and k 2D convolution results can be obtained. Since the activation data for successfully recognizing the data to be recognized is carried in these convolution results, the pose type recognized based on the activation data at multiple levels can more accurately reflect the actual pose type of the target object.
[0251] Exemplarily, when recognizing the pose type, the pose type with the highest probability can be selected from multiple different pose types as the final recognition result.
[0252] It should be noted that the value range of k can be determined by training the neural network that implements the recognition method provided in the embodiments of the present application. Exemplarily, the value range of k can be [0, K]. K can be an integer greater than 1.
[0253] As described above, in the recognition method provided in the embodiments of the present application, activation data at multiple levels are obtained through processing at multiple levels, thereby achieving full activation of each local feature in the data to be recognized, so that even when some bone nodes are occluded or the network fails, the pose type of the target object can still be accurately recognized.
[0254] From the above, it can be seen that in the recognition method provided in the embodiments of the present application, after obtaining the data to be recognized including at least two pose sequences, first feature extraction is performed on the data to be recognized to obtain the first-level graph convolution data, thereby reducing the redundant data or interference data carried by the first graph convolution data and laying a foundation for subsequent pose type recognition; then, the first-level three-dimensional graph convolution processing and the first-level two-dimensional graph convolution processing are respectively performed on the first-level graph convolution data to obtain the first-level three-dimensional convolution result including the local feature information of the data to be recognized and the first-level two-dimensional convolution result including the association relationship between each local feature information, and then the first-level activation data is determined according to the first-level three-dimensional convolution result and the first-level two-dimensional convolution result, so that the first-level activation data contains comprehensive features of the data to be recognized; finally, the pose type is recognized based on the first-level activation data and the first-level graph convolution data, so that even when the data acquisition device is occluded or stuck, that is, the data to be recognized does not carry complete feature data of the target object, the recognition method provided in the present application can still resist the risk brought by missing information based on each local feature information and the association relationship between the local feature information, and accurately recognize the data to be recognized, thereby achieving efficient and accurate recognition of the pose type in a wider range of application scenarios.
[0255] Based on the foregoing embodiments, the embodiments of the present application provide a neural network training method. Figure 7 It is a schematic flowchart of the first neural network training method provided in the embodiments of the present application.
[0256] This neural network training method can be implemented by a processor of a neural network training device. The above-mentioned processor can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor.
[0257] Exemplarily, the neural network can include a first network and a second network.
[0258] Among them, the first network is used for feature extraction; the second network is used for graph convolution processing on the output data of the first network.
[0259] The second network includes at least a first-level three-dimensional graph convolutional network for implementing first-level three-dimensional graph convolutional processing and a first-level two-dimensional graph convolutional network for implementing first-level two-dimensional graph convolutional processing.
[0260] In one embodiment, the first network can be a linear network or a non-linear network.
[0261] In one embodiment, the first network can be a perceptron. Exemplarily, the perceptron can be multi-layer.
[0262] In one embodiment, the first-level three-dimensional graph convolutional network and the first-level three-dimensional graph convolutional network in the second network can be any neural network capable of performing convolutional calculations on graph data.
[0263] Figure 7 The neural network training method shown can include the following steps:
[0264] Step 701, obtain sample data including multiple pose sequences.
[0265] Wherein, the pose sequence includes multiple node information of any pose of any object.
[0266] In one embodiment, the sample data can include pose sequences of multiple objects of multiple types.
[0267] In one embodiment, the sample data can include pose sequences of objects of a certain specified type.
[0268] In one embodiment, the sample data can include pose sequences of objects of a specified type in a specified state. For example, the pose sequence of the first object when it is in a running state.
[0269] In one embodiment, the sample data can include pose sequences of any object in a specified scenario. Exemplarily, the specified scenario can include at least one scenario; the specified scenario can include smart home, security monitoring, education and entertainment, etc.
[0270] In one embodiment, the sample data can be data carrying pose type labels. Through this labeled data, the pose type of any object obtained after processing the sample data by the neural network can be used to adjust each parameter in the neural network.
[0271] In one embodiment, the sample data can be stored in a database.
[0272] Step 702, based on the first network, perform feature extraction on the sample data to obtain a first-level feature sequence.
[0273] Among them, the first-level feature sequence includes the feature sequences of each pose sequence.
[0274] In one implementation, feature extraction of the sample data can be linear feature extraction of the sample data through the first network.
[0275] In one implementation, feature extraction of the sample data can be non-linear feature extraction of the sample data through the first network.
[0276] In one implementation, feature extraction of the sample data can be to determine the feature extraction parameters of the first network according to the attribute parameters of the sample data, and then perform feature extraction on the sample data based on the feature extraction parameters.
[0277] In one implementation, feature extraction of the sample data can be achieved through Equation (5) in the foregoing embodiment under the condition that the input data is the sample data.
[0278] Step 703: Based on the first-level three-dimensional graph convolutional network, perform first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level feature sequence to obtain a first-level three-dimensional result.
[0279] Among them, the first-level three-dimensional result includes the local feature information in the first-level feature sequence.
[0280] Exemplarily, the process of obtaining the first-level three-dimensional result is the same as the process of obtaining the first-level three-dimensional convolution result, which will not be elaborated here.
[0281] Step 704: Based on the first-level two-dimensional graph convolutional network, perform first-level two-dimensional graph convolution processing on the first-level feature sequence to obtain a first-level two-dimensional result.
[0282] Among them, the first-level two-dimensional result includes the correlation relationship between the local feature information in the first feature sequence.
[0283] It should be noted that Step 703 and Step 704 can be executed simultaneously, and the order can also be adjusted successively. The embodiments of the present application do not make any limitations in this regard.
[0284] Step 705: Based on the first-level feature sequence, the first-level three-dimensional result, and the first-level two-dimensional result, identify the pose type of any object.
[0285] Step 706: Based on the pose type, train the first network and the second network to obtain a trained neural network.
[0286] Thus, in the neural network training method provided in the embodiment of the present application, the first-level feature sequence corresponding to each posture sequence in the sample data is first obtained through the first network, and then the first-level three-dimensional graph convolution processing is performed on at least two feature sequences in the first-level feature sequence to obtain a first-level three-dimensional result containing local feature information in the first-level graph convolution data, and the first-level two-dimensional graph convolution processing is performed on the first-level feature sequence to obtain a first-level two-dimensional result, and then based on the first-level feature sequence, the first-level three-dimensional result and the first-level two-dimensional result, the posture type of any object is identified, and then the first network and the second network are trained according to the posture type to obtain a trained neural network.
[0287] From the above, it can be seen that in the neural network training method provided in the embodiment of the present application, the first-level three-dimensional result carries the local feature information in the first-level feature sequence, and the first-level two-dimensional result carries the association relationship between the local feature information in the first-level feature sequence. On this basis, through the combination of the first-level three-dimensional result and the first-level two-dimensional result, it is possible to obtain a set of local feature information with established associations carried by the sample data, that is, to achieve the recognition of each dimension of the features carried by the sample data, and to construct a deeper feature recognition process for the sample data. Therefore, even in the case where the data acquisition device is blocked or stuck, that is, the data to be identified does not carry the complete feature data of the target object, the neural network obtained by the neural network training method provided by the present application can also be based on the local feature information and the association relationship between the local feature information. Accurately identify the data to be identified, resist the risks caused by missing information, and thus achieve efficient and accurate recognition of posture types in a wider range of application scenarios.
[0288] Based on the above embodiments, the present application provides a second neural network training method. Figure 8 A schematic diagram of a flow chart of a second neural network training method provided in an embodiment of the present application.
[0289] The neural network includes a first network and a second network. The first network is used for feature extraction; the second network is used for performing graph convolution processing on the output data of the first network; the second network at least includes a first-level three-dimensional graph convolution network for implementing a first-level three-dimensional graph convolution processing, and a first-level two-dimensional graph convolution network for implementing a first-level two-dimensional graph convolution processing.
[0290] The flowchart of the neural network training method may include the following steps:
[0291] Step 801: Obtain sample data including multiple posture sequences.
[0292] The posture sequence includes multiple node information of any posture of any object.
[0293] Step 802: Based on the first network, extract features from the sample data to obtain a first-level feature sequence.
[0294] Among them, the first-level feature sequence includes the feature sequences of each pose sequence.
[0295] Exemplarily, step 802 can be implemented through steps E1 - E3:
[0296] Step E1: Obtain third data from the sample data.
[0297] Among them, the third data represents data of at least two dimensions corresponding to the sample data.
[0298] Exemplarily, the way to obtain the third data is the same as the way to obtain the first data in the foregoing embodiment, that is, it can be obtained by processing the sample data through formulas (1) - (3).
[0299] Step E2: Based on the first network, respectively extract features from the data of each dimension in the third data to obtain fourth data.
[0300] Exemplarily, the way to obtain the fourth data is the same as the way to obtain the second data in the foregoing embodiment, that is, it can be obtained by successively processing the sample data and the third data through formula (5) and Figure 4 the manner shown.
[0301] Step E3: Based on the first network, process the fourth data to obtain a first-level feature sequence.
[0302] Exemplarily, the process of obtaining the first-level feature sequence is the same as the process of obtaining the first-level graph convolutional data in the foregoing embodiment, and will not be elaborated here.
[0303] Step 803: Based on the first-level three-dimensional graph convolutional network, perform first-level three-dimensional graph convolutional processing on at least two feature sequences in the first-level feature sequence to obtain a first-level three-dimensional result.
[0304] Among them, the first-level three-dimensional result includes local feature information in the first-level feature sequence.
[0305] Exemplarily, step 803 can be implemented through steps F1 - step F4:
[0306] Step F1: Determine a sliding window.
[0307] Step F2: Based on the sliding window, obtain at least two feature sequences from the first-level feature sequence.
[0308] Exemplarily, based on a sliding window, at least two feature sequences are obtained from the first-level feature sequence, which can be implemented by the same method as step B2 in the foregoing embodiment, and will not be elaborated here.
[0309] Step F3: Determine a first-level three-dimensional matrix based on at least two feature sequences.
[0310] The first-level three-dimensional matrix includes each node in at least two feature sequences and the association relationship between any two nodes in at least two feature sequences.
[0311] Exemplarily, the determination process of the first-level three-dimensional matrix can be the same as the determination process of the first-level three-dimensional data in the foregoing embodiment, that is, at least two feature sequences in the first-level feature sequence are processed by equations (6)-(7).
[0312] Step F4: Perform first-level three-dimensional graph convolution processing on the first-level three-dimensional matrix based on the first-level three-dimensional graph convolution network.
[0313] Exemplarily, the first-level three-dimensional graph convolution processing can process the first-level three-dimensional matrix through the same process as the first-level three-dimensional graph convolution processing in the foregoing embodiment, that is, process the first-level three-dimensional matrix by equation (8).
[0314] Step 804: Perform first-level two-dimensional graph convolution processing on the first-level feature sequence based on the first-level two-dimensional graph convolution network to obtain a first-level two-dimensional result.
[0315] The first-level two-dimensional result includes the association relationship between local feature information in the first feature sequence.
[0316] It should be noted that step 803 and step 804 can be executed simultaneously or in reverse order, and the embodiments of the present application do not limit this.
[0317] Step 805: Fuse the first-level three-dimensional result and the first-level two-dimensional result to obtain a first-level recognition result.
[0318] Exemplarily, the process of obtaining the first-level recognition result can be the same as the process of obtaining the first-level activation data.
[0319] Step 806: Identify the pose type of any object based on the first-level feature sequence and the first-level recognition result.
[0320] Exemplarily, the second network further includes a kth-level three-dimensional graph convolution network for implementing the kth-level three-dimensional graph convolution processing and a kth-level two-dimensional graph convolution network for implementing the kth-level two-dimensional graph convolution processing; where k is an integer greater than 1.
[0321] Step 806 can be implemented through Steps G1 - G3:
[0322] Step G1: Based on the (k - 1)-th level recognition result and the (k - 1)-th level feature sequence, determine the k-th level feature sequence.
[0323] Among them, the k-th level feature sequence includes the feature sequence that was not recognized in the (k - 1)-th level feature sequence.
[0324] Exemplarily, the k-th level feature sequence can be determined by the same method as the k-th level convolution data in the foregoing embodiment.
[0325] Step G2: Process the k-th level feature sequence through the k-th level three-dimensional graph convolution network and the k-th level two-dimensional graph convolution network to obtain the k-th level recognition result.
[0326] Among them, the k-th level recognition result represents the feature information recognized from the k-th level feature sequence.
[0327] Exemplarily, Step G2 can be implemented through Steps H1 - H3:
[0328] Step H1: Based on the k-th level three-dimensional graph convolution network, perform k-th level three-dimensional graph convolution processing on at least two feature sequences in the k-th level feature sequence to obtain the k-th level three-dimensional convolution result.
[0329] Exemplarily, the k-th level three-dimensional convolution result can be obtained by processing at least two feature sequences in the k-th level feature sequence through the method shown in the k-th level three-dimensional graph convolution processing provided in the foregoing embodiment, that is, by processing at least two feature sequences in the k-th level feature sequence through Equation (12).
[0330] Step H2: Based on the k-th level two-dimensional graph convolution network, perform k-th level two-dimensional graph convolution processing on the k-th level feature sequence to obtain the k-th level two-dimensional convolution result.
[0331] Exemplarily, the process of obtaining the k-th level two-dimensional convolution result is the same as that of the k-th level two-dimensional convolution result provided in the foregoing embodiment, and will not be elaborated here.
[0332] Step H3: Based on the k-th level two-dimensional convolution result and the k-th level three-dimensional convolution result, obtain the k-th level recognition result.
[0333] Exemplarily, the process of obtaining the k-th level recognition result can be the same as that of the k-th level activation data provided in the foregoing embodiment, and will not be elaborated here.
[0334] Step G3: Based on the first level recognition result to the k-th level recognition result, recognize the pose type of any object.
[0335] Step 807: Based on the pose type, train the first network and the second network to obtain a trained neural network.
[0336] Exemplarily, step 807 can be implemented through steps J1 - J3:
[0337] Step J1: Obtain an error threshold and an expected recognition result.
[0338] In one implementation, the error threshold can be preset.
[0339] In one implementation, the error threshold can vary according to the type of sample data.
[0340] In one implementation, the error threshold can be adjusted according to the different application scenarios corresponding to the sample data.
[0341] In one implementation, the error threshold can be adjusted according to the type of any object.
[0342] In one implementation, the expected recognition result can represent the true pose type corresponding to the sample data.
[0343] In one implementation, the expected recognition result can be obtained from the label data carried by the sample data.
[0344] Step J2: Based on the first - level recognition result to the k - th level recognition result and the expected recognition result, determine the recognition error.
[0345] In one implementation, the recognition error can be achieved in the following way:
[0346] Determine the neural network recognition result through the first - level recognition result to the k - th level recognition result;
[0347] Determine the recognition error according to the neural network recognition result and the expected recognition result.
[0348] Exemplarily, the neural network recognition result can be obtained by weighted superposition of the first - level recognition result to the k - th level recognition result.
[0349] In one implementation, the recognition error can be measured by the recognition loss value of the neural network.
[0350] Exemplarily, the recognition loss value of the neural network can be determined by the loss function shown in Equation (13):
[0351]
[0352] In Equation (13), For representing the recognition result of the neural network; p k For representing the expected recognition result; L loss For representing the recognition loss value of the neural network.
[0353] Step J3: Based on the shown error threshold and the recognition error, train the first network and the second network through the backpropagation algorithm to obtain a trained neural network.
[0354] In one implementation, when the recognition error is less than or equal to the error threshold, it indicates that the training result of the neural network can already efficiently recognize the pose information of any object. At this time, the training of the neural network can be stopped.
[0355] In one implementation, when the recognition error is greater than the error threshold, the first network and the second network can be trained through the backpropagation algorithm to obtain a trained neural network.
[0356] Exemplarily, the above training process of the neural network can be performed by randomly selecting any sample data from the sample data, and the above training process of the neural network is repeatedly executed until the recognition error is less than or equal to the error threshold. At this time, the determined k is the number of levels K of the finally determined three-dimensional graph convolutional network and two-dimensional graph convolutional network in the neural network.
[0357] As can be seen from the above, the neural network training method provided in the embodiments of the present application, after obtaining sample data including at least two pose sequences, performs first feature extraction on the sample data to obtain a first-level feature sequence, so as to reduce the redundant data or interference data carried by the first-level feature sequence and lay a foundation for subsequent pose type recognition; then perform first-level three-dimensional graph convolution processing and first-level two-dimensional graph convolution processing on the first-level feature sequence respectively to obtain a first-level three-dimensional result including local feature information of the data to be recognized and a first-level two-dimensional result including the association relationship between each local feature information, and then determine the first-level recognition result according to the first-level three-dimensional result and the first-level two-dimensional result, so that the first-level recognition result contains comprehensive features of the sample data; finally, based on the first-level recognition result and the first-level feature sequence, recognize the pose type, so that even when the data acquisition device is blocked or stuck, that is, the data to be recognized does not carry complete feature data of the target object, the neural network obtained by the neural network training method provided in the present application can also accurately recognize the data to be recognized based on each local feature information and the association relationship between the local feature information, thus realizing efficient and accurate recognition of pose types in a wider application scenario.
[0358] Based on the foregoing embodiments, the embodiments of the present application provide an identification device 9, Figure 9Schematic diagram of the recognition device 9 provided by the embodiments of the present application.
[0359] The recognition device 9 includes a first acquisition module 901, a first processing module 902, and a first recognition module 903. Among them:
[0360] The first acquisition module 901 is configured to acquire the data to be recognized including at least two pose sequences; wherein, the pose sequence includes multiple node information of any pose of the target object;
[0361] The first processing module 902 is configured to perform feature extraction on the data to be recognized to obtain first-level graph convolution data; perform first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level graph convolution data to obtain a first-level three-dimensional convolution result; perform first-level two-dimensional graph convolution processing on the first-level graph convolution data to obtain a first-level two-dimensional convolution result; wherein, the first-level graph convolution data includes the feature sequences of each pose sequence; the first-level three-dimensional convolution result includes the local feature information in the first-level graph convolution data; the first-level two-dimensional convolution result includes the correlation relationship between the local feature information in the first-level graph convolution data;
[0362] The first recognition module 903 is configured to recognize the pose type of the target object based on the first-level graph convolution data, the first-level three-dimensional convolution result, and the first-level two-dimensional convolution result.
[0363] In some embodiments, the first processing module 902 is configured to fuse the first-level three-dimensional convolution result and the first-level two-dimensional convolution result to obtain first-level activation data; recognize the pose type based on the first-level activation data and the first-level graph convolution data; wherein, the first-level activation data represents the feature information recognized from the first-level graph convolution data.
[0364] In some embodiments, the first processing module 902 is configured to determine the k-level graph convolution data based on the (k - 1)-level activation data and the (k - 1)-level graph convolution data; where k is an integer greater than 1; obtain the k-level activation data based on the k-level graph convolution data; the k-level graph convolution data includes the feature sequences in the (k - 1)-level graph convolution data that have not been recognized; wherein, the k-level activation data represents the feature information recognized from the k-level graph convolution data.
[0365] The first recognition module 903 is configured to recognize the pose type based on the first-level activation data to the k-level activation data;
[0366] In some embodiments, the first processing module 902 is configured to perform a k-level three-dimensional graph convolution process on at least two feature sequences in the k-level graph convolution data to obtain a k-level three-dimensional convolution result; perform a k-level two-dimensional graph convolution process on the k-level graph convolution data to obtain a k-level two-dimensional convolution result; and obtain a k-level activation data based on the k-level three-dimensional convolution result and the k-level two-dimensional convolution result.
[0367] In some embodiments, the first processing module 902 is configured to determine a sliding window; obtain at least two feature sequences from the first-level graph convolution data based on the sliding window; determine first-level three-dimensional data based on the at least two feature sequences; and perform a first-level three-dimensional graph convolution process on the first-level three-dimensional data; wherein the first-level three-dimensional data includes each node in the at least two feature sequences and the association relationship between any two nodes in the at least two feature sequences.
[0368] In some embodiments, the first processing module 902 is configured to obtain first data from the data to be recognized; perform feature extraction on each dimension of the first data respectively to obtain second data; and obtain first-level graph convolution data based on the second data; wherein the first data represents at least two-dimensional data corresponding to the data to be recognized.
[0369] In some embodiments, the at least two dimensions include at least two of the following dimensions: distance dimension, speed dimension, and position dimension.
[0370] As can be seen from the above, in the first-level three-dimensional graph convolution result obtained by the recognition device provided in the embodiments of the present application, the local feature information in the first-level graph convolution data is carried, and in the first-level two-dimensional graph convolution result, the association relationship between the local feature information in the first-level graph convolution data is carried. On this basis, by combining the first-level three-dimensional convolution result and the first-level two-dimensional convolution result, a set of local feature information with associated relationships carried by the data to be recognized can be obtained, that is, the recognition of each dimension of the features carried by the data to be recognized is realized, and a deeper feature recognition process for the data to be recognized is constructed. In this way, even when the data acquisition device is blocked or stuck, that is, the data to be recognized does not carry complete feature data of the target object, the recognition device provided in the present application can still accurately recognize the data to be recognized based on each local feature information and the association relationship between the local feature information, so as to realize efficient and accurate recognition of the posture type in a wider range of application scenarios.
[0371] Based on the foregoing embodiments, an embodiment of the present application provides a neural network training device 10. Figure 10 It is a schematic structural diagram of the neural network training device 10 provided in the embodiment of the present application.
[0372] The neural network includes a first network and a second network; among them, the first network is used for feature extraction; the second network is used for performing graph convolution processing on the output data of the first network; the second network at least includes a first-level three-dimensional graph convolution network for implementing first-level three-dimensional graph convolution processing and a first-level two-dimensional graph convolution network for implementing first-level two-dimensional graph convolution processing; the neural network training device 10 includes a second acquisition module 1001, a second processing module 1002, and a second recognition module 1003; where:
[0373] The second acquisition module 1001 is used to acquire sample data including multiple pose sequences; among them, the pose sequence includes multiple node information of any pose of any object;
[0374] The second processing module 1002 is used to perform feature extraction on the sample data based on the first network to obtain a first-level feature sequence; perform first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level feature sequence based on the first-level three-dimensional graph convolution network to obtain a first-level three-dimensional result; perform first-level two-dimensional graph convolution processing on the first-level feature sequence based on the first-level two-dimensional graph convolution network to obtain a first-level two-dimensional result; where the first-level feature sequence includes the feature sequence of each pose sequence; the first-level three-dimensional result includes local feature information in the first-level feature sequence; the first-level two-dimensional result includes the correlation relationship between the local feature information in the first feature sequence;
[0375] The second recognition module 1003 is used to recognize the pose type of any object based on the first-level feature sequence, the first-level three-dimensional result, and the first-level two-dimensional result;
[0376] The second processing module 1002 is further used to train the first network and the second network based on the pose type to obtain a trained neural network.
[0377] In some embodiments, the second processing module 1002 is used to fuse the first-level three-dimensional result and the first-level two-dimensional result to obtain a first-level recognition result; recognize the pose type of any object based on the first-level feature sequence and the first-level recognition result.
[0378] In some embodiments, the second network further includes a k-level three-dimensional graph convolution network for implementing k-level three-dimensional graph convolution processing and a k-level two-dimensional graph convolution network for implementing k-level two-dimensional graph convolution processing; where k is an integer greater than 1;
[0379] The second processing module 1002 is configured to determine the k-th level feature sequence based on the (k - 1)-th level recognition result and the (k - 1)-th level feature sequence; process the k-th level feature sequence through the k-th level 3D graph convolutional network and the k-th level 2D graph convolutional network to obtain the k-th level recognition result; recognize the pose type of any object based on the first level recognition result to the k-th level recognition result; the k-th level feature sequence includes the feature sequences in the (k - 1)-th level feature sequence that have not been recognized; wherein, the k-th level recognition result represents the feature information recognized from the k-th level feature sequence.
[0380] In some embodiments, the second processing module 1002 is configured to perform k-th level 3D graph convolution processing on at least two feature sequences in the k-th level feature sequence based on the k-th level 3D graph convolutional network to obtain the k-th level 3D convolution result;
[0381] The second processing module 1002 is further configured to perform k-th level 2D graph convolution processing on the k-th level feature sequence based on the k-th level 2D graph convolutional network to obtain the k-th level 2D convolution result; obtain the k-th level recognition result based on the k-th level 2D convolution result and the k-th level 3D convolution result.
[0382] In some embodiments, the second acquisition module 1001 is configured to acquire an error threshold and an expected recognition result;
[0383] The second processing module 1002 is configured to determine the recognition error based on the first level recognition result to the k-th level recognition result and the expected recognition result; train the first network and the second network through the backpropagation algorithm based on the error threshold and the recognition error to obtain the trained neural network.
[0384] In some embodiments, the second acquisition module 1001 is configured to determine a sliding window;
[0385] The second processing module 1002 is configured to acquire at least two feature sequences from the first level feature sequence based on the sliding window; determine the first level 3D matrix based on the at least two feature sequences; wherein, the first level 3D matrix includes each node in the at least two feature sequences and the association relationship between any two nodes in the at least two feature sequences;
[0386] The second processing module 1002 is further configured to perform first level 3D graph convolution processing on the first level 3D matrix based on the first level 3D graph convolutional network.
[0387] In some embodiments, the second acquisition module 1001 is configured to acquire third data from the sample data; wherein, the third data represents data of at least two dimensions corresponding to the sample data;
[0388] The second processing module 1002 is configured to perform feature extraction on each dimension of the third data based on the first network to obtain fourth data; and process the fourth data based on the first network to obtain a first-level feature sequence.
[0389] As can be seen from the above, the first-level three-dimensional result obtained by the neural network training device provided by the embodiment of the present application carries local feature information in the first-level feature sequence; the first-level two-dimensional result carries the correlation relationship between the local feature information in the first-level feature sequence. On this basis, by combining the first-level three-dimensional result and the first-level two-dimensional result, a set of local feature information with established correlation relationships carried by the sample data can be obtained, that is, the recognition of each dimension of the features carried by the sample data is realized, and a deeper feature recognition process of the sample data is constructed. Even when the data acquisition device is blocked or stuck, that is, the data to be recognized does not carry complete feature data of the target object, the neural network trained by the neural network training device provided by the present application can also accurately recognize the data to be recognized based on each local feature information and the correlation relationship between the local feature information, thereby realizing efficient and accurate recognition of posture types in a wider range of application scenarios.
[0390] Based on the foregoing embodiments, an embodiment of the present application provides a behavior recognition system 11 based on a bone sequence. Figure 11 It is a schematic structural diagram of the behavior recognition system 11 provided by the embodiment of the present application. As Figure 11 shown, the behavior recognition system 11 may include an identification device 9 and a neural network training device 10, where:
[0391] In Figure 11 it, the identification device 9 can perform state recognition on various bone sequences based on the neural network trained by the neural network training device 10.
[0392] Among them, various bone sequences processed by the identification device 9 can be obtained by processing the posture data collected by a millimeter-wave radar, an ordinary camera, and the data collected by an infrared camera. Exemplarily, an infrared camera or an ordinary camera can be configured with OpenPose or the like to implement the posture estimation of the target object.
[0393] The processing of the bone sequence is realized through a generalized graph convolutional behavior recognition module.
[0394] Exemplarily, Figure 11 the generalized graph convolutional behavior recognition module in
[0395] By processing the bone sequence through the generalized graph convolution behavior recognition module, the activation data from the first level to the Kth level can be obtained, and based on the activation data from the first level to the Kth level, the pose type can be recognized.
[0396] After obtaining the pose type, it can also be input to the display device through the display device to display the recognition result of the recognition device 9 in a more intuitive way.
[0397] In Figure 11 the neural network used for bone sequence recognition in the recognition device 9 is trained based on the sample data by the Figure 10 neural network training device 10 in
[0398] In Figure 11 the neural network training device 10 shown, first, input information encoding is performed. This operation can extract features from the input sample data through the first network in any previous embodiment. Exemplarily, the sample data can be feature-extracted through a multi-layer perceptron to obtain the first-level feature sequence.
[0399] Then, the first-level generalized graph convolution is performed. Exemplarily, the first-level generalized graph convolution operation can be implemented by processing the first-level feature sequence through the first-level three-dimensional graph convolution network and the first-level two-dimensional graph convolution network in the foregoing embodiments.
[0400] Secondly, the first-level activation mapping operation is performed. Exemplarily, after the first-level generalized graph convolution operation, the first-level three-dimensional result and the first-level two-dimensional result can be obtained, and the first-level recognition result can be obtained based on the first-level three-dimensional result and the first-level two-dimensional result.
[0401] Exemplarily, for the second-level generalized graph convolution operation, after the first-level recognition result and the first-level feature sequence are subjected to the first-level activation mapping, the bone features not recognized by the first-level generalized graph convolution can be obtained; and the bone features not recognized by the first-level generalized graph convolution are subjected to the second-level generalized graph convolution operation.
[0402] Exemplarily, the second-level generalized graph convolution operation can be implemented by processing the bone features not recognized by the first-level generalized graph convolution through the second-level three-dimensional graph convolution network and the second-level two-dimensional graph convolution network in the foregoing embodiments.
[0403] Exemplarily, the second-level activation mapping operation can be performed after the second-level generalized graph convolution is performed on the bone features not recognized by the first-level generalized graph convolution to obtain the second-level recognition result. Through the second-level activation mapping operation, the bone features not recognized by the second-level generalized graph convolution can be obtained, and the bone features not recognized by the second-level generalized graph convolution are subjected to the kth-level generalized graph convolution to obtain the kth-level recognition result.
[0404] Subsequently, the predicted behavior probability is obtained through the skeletal features. Exemplarily, according to the skeletal features carried in the recognition results from the first level to the Kth level, the predicted behavior probability can be obtained.
[0405] Calculate the loss value of the neural network. Exemplarily, the loss value of the neural network can be calculated based on the predicted behavior probability and the true behavior.
[0406] Then, the backpropagation algorithm is used to update the network parameters. Exemplarily, the updated network parameters include any network parameters of the first network and the second network as described above, or can also be the parameter update of a specified network or a sub-network.
[0407] Finally, based on random samples, the above steps are repeatedly executed until the training ends. Exemplarily, the end of training can be determined by the recognition error and the error threshold.
[0408] In this way, the behavior recognition system 11 based on the skeletal sequence provided by the embodiments of the present application can, on the basis of the neural network trained based on random samples, when recognizing the skeletal sequence, respectively extract the local features corresponding to the skeletal sequence and the correlation relationship between the local features. Therefore, even when the skeletal data is partially occluded or the data acquisition device freezes, it can accurately identify the pose type corresponding to the skeletal sequence according to the local features and the correlation relationship between the local features, thereby generalizing the pose type recognition based on the skeletal sequence.
[0409] Based on the foregoing embodiments, an electronic device 12 is provided in the embodiments of the present application. Figure 12 It is a schematic structural diagram of the electronic device 12 provided in the embodiments of the present application.
[0410] As Figure 12 shown, the electronic device 12 may include a processor 1201, a memory 1202, and a communication bus, where:
[0411] The communication bus is used to implement the communication connection between the processor 1201 and the memory 1202; the processor 1201 is used to execute the computer program stored in the memory 1202 to implement the recognition method of any previous embodiment or the neural network training method of any previous embodiment.
[0412] Among them, the above-mentioned processor 1201 may be at least one of an application-specific integrated circuit ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor. It can be understood that other electronic devices for implementing the above processor functions may also be used, and the embodiments of the present invention do not make specific limitations.
[0413] The above-mentioned memory 1202 can be a volatile memory, such as RAM; or a non-volatile memory, such as ROM, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memories, and provides instructions and data to the processor.
[0414] In some embodiments, the functions or modules included in the device provided by the embodiments of the present application can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0415] The descriptions of the above embodiments tend to emphasize the differences between the embodiments. The similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated in this article.
[0416] The methods disclosed in the method embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments.
[0417] The features disclosed in the product embodiments provided by the present application can be arbitrarily combined without conflict to obtain new product embodiments.
[0418] The features disclosed in the method or device embodiments provided by the present application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0419] It should be noted that the above computer-readable storage medium may be a read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; it may also be various electronic devices including one or any combination of the above memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0420] It should be noted that in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including that element.
[0421] The serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.
[0422] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0423] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0424] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0425] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0426] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A recognition method, characterized in that, The method includes: Obtaining the data to be recognized including at least two pose sequences; wherein, each pose sequence includes multiple node information of any pose of the target object; Performing feature extraction on the data to be recognized to obtain first-level graph convolutional data; wherein, the first-level graph convolutional data includes a feature sequence of each pose sequence; Performing first-level three-dimensional graph convolution processing on at least two of the feature sequences in the first-level graph convolutional data to obtain a first-level three-dimensional convolution result; wherein, the first-level three-dimensional convolution result includes local feature information in the first-level graph convolutional data; Performing first-level two-dimensional graph convolution processing on the first-level graph convolutional data to obtain a first-level two-dimensional convolution result; wherein, the first-level two-dimensional convolution result includes the correlation relationship between local feature information in the first-level graph convolutional data; Identifying the pose type of the target object based on the first-level graph convolutional data, the first-level three-dimensional convolution result, and the first-level two-dimensional convolution result.
2. The method according to claim 1, wherein The identifying the pose type of the target object based on the first-level graph convolutional data, the first-level three-dimensional convolution result, and the first-level two-dimensional convolution result includes: Fusing the first-level three-dimensional convolution result and the first-level two-dimensional convolution result to obtain first-level activation data; wherein, the first-level activation data represents the feature information identified from the first-level graph convolutional data; Identifying the pose type based on the first-level activation data and the first-level graph convolutional data.
3. The method according to claim 2, characterized in that, The identifying the pose type based on the first-level activation data and the first-level graph convolutional data includes: Determining k-th level graph convolutional data based on the (k - 1)-th level activation data and the (k - 1)-th level graph convolutional data; where k is an integer greater than 1; the k-th level graph convolutional data includes the feature sequences in the (k - 1)-th level graph convolutional data that have not been recognized; Obtaining k-th level activation data based on the k-th level graph convolutional data; wherein, the k-th level activation data represents the feature information identified from the k-th level graph convolutional data; Identifying the pose type based on the first-level activation data to the k-th level activation data.
4. The method according to claim 3, wherein The obtaining the k-th level activation data based on the k-th level graph convolutional data includes: Performing k-th level three-dimensional graph convolution processing on at least two feature sequences in the k-th level graph convolutional data to obtain a k-th level three-dimensional convolution result; Performing k-th level two-dimensional graph convolution processing on the k-th level graph convolutional data to obtain a k-th level two-dimensional convolution result; Obtaining the k-th level activation data based on the k-th level three-dimensional convolution result and the k-th level two-dimensional convolution result.
5. The method according to claim 1, characterized in that, The performing first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level graph convolutional data includes: Determining a sliding window; Based on the sliding window, obtaining at least two of the feature sequences from the first-level graph convolutional data; Determine the first-level three-dimensional data based on at least two of the feature sequences; wherein, the first-level three-dimensional data includes each node in at least two of the feature sequences, and the association relationship between any two nodes in at least two of the feature sequences; Perform the first-level three-dimensional graph convolution processing on the first-level three-dimensional data.
6. The method according to claim 1, wherein The feature extraction of the data to be recognized to obtain the first-level graph convolution data includes: Obtain the first data from the data to be recognized; wherein, the first data represents data corresponding to at least two dimensions of the data to be recognized; Perform feature extraction on the data of each dimension in the first data respectively to obtain the second data; Based on the second data, obtain the first-level graph convolution data.
7. The method according to claim 6, characterized in that, The at least two dimensions include at least two of the following dimensions: distance dimension, speed dimension, position dimension.
8. A neural network training method, characterized in that, The neural network includes a first network and a second network; wherein, the first network is used for feature extraction; the second network is used for performing graph convolution processing on the output data of the first network; the second network at least includes a first-level three-dimensional graph convolution network for implementing the first-level three-dimensional graph convolution processing and a first-level two-dimensional graph convolution network for implementing the first-level two-dimensional graph convolution processing; the method includes: Obtain sample data including multiple pose sequences; wherein, the pose sequence includes multiple node information of any pose of any object; Based on the first network, perform feature extraction on the sample data to obtain the first-level feature sequences; wherein, the first-level feature sequences include the feature sequences of each of the pose sequences; Based on the first-level three-dimensional graph convolution network, perform the first-level three-dimensional graph convolution processing on at least two of the first-level feature sequences to obtain the first-level three-dimensional result; wherein, the first-level three-dimensional result includes the local feature information in the first-level feature sequences; Based on the first-level two-dimensional graph convolution network, perform the first-level two-dimensional graph convolution processing on the first-level feature sequences to obtain the first-level two-dimensional result; wherein, the first-level two-dimensional result includes the association relationship between the local feature information in the first-level feature sequences; Based on the first-level feature sequences, the first-level three-dimensional result and the first-level two-dimensional result, identify the pose type of the any object; Based on the pose type, train the first network and the second network to obtain the trained neural network.
9. The method according to claim 8, wherein Based on the first-level feature sequences, the first-level three-dimensional result and the first-level two-dimensional result, identifying the pose type of the any object includes: Fuse the first-level three-dimensional result and the first-level two-dimensional result to obtain the first-level recognition result; Based on the first-level feature sequences and the first-level recognition result, identify the pose type.
10. The method according to claim 9, characterized in that, The second network further includes a k-th level three-dimensional graph convolution network for implementing the k-th level three-dimensional graph convolution processing and a k-th level two-dimensional graph convolution network for implementing the k-th level two-dimensional graph convolution processing; where k is an integer greater than 1; the identifying the pose type based on the first-level feature sequence and the first-level recognition result includes: Determining a k-th level feature sequence based on the (k - 1)-th level recognition result and the (k - 1)-th level feature sequence; where the k-th level feature sequence includes the feature sequences in the (k - 1)-th level feature sequence that have not been recognized; Processing the k-th level feature sequence through the k-th level three-dimensional graph convolution network and the k-th level two-dimensional graph convolution network to obtain a k-th level recognition result; where the k-th level recognition result represents the feature information recognized from the k-th level feature sequence; Identifying the pose type based on the first-level recognition result to the k-th level recognition result.
11. The method according to claim 10, characterized in that, The processing the k-th level feature sequence through the k-th level three-dimensional graph convolution network and the k-th level two-dimensional graph convolution network to obtain a k-th level recognition result includes: Performing the k-th level three-dimensional graph convolution processing on at least two feature sequences in the k-th level feature sequence based on the k-th level three-dimensional graph convolution network to obtain a k-th level three-dimensional convolution result; Performing the k-th level two-dimensional graph convolution processing on the k-th level feature sequence based on the k-th level two-dimensional graph convolution network to obtain a k-th level two-dimensional convolution result; Obtaining the k-th level recognition result based on the k-th level two-dimensional convolution result and the k-th level three-dimensional convolution result.
12. The method according to claim 10, characterized in that, Training the first network and the second network based on the pose type to obtain a trained neural network, including: Obtaining an error threshold and an expected recognition result; Determining a recognition error based on the first-level recognition result to the k-th level recognition result and the expected recognition result; Training the first network and the second network through a backpropagation algorithm based on the error threshold and the recognition error to obtain the trained neural network.
13. The method according to claim 8, characterized in that, The performing the first-level three-dimensional graph convolution processing on at least two feature sequences in the first-level feature sequence based on the first-level three-dimensional graph convolution network includes: Determining a sliding window; Obtaining at least two of the feature sequences from the first-level feature sequence based on the sliding window; Determining a first-level three-dimensional matrix based on at least two of the feature sequences; where the first-level three-dimensional matrix includes each node in at least two of the feature sequences and the association relationship between any two nodes in at least two of the feature sequences; Performing the first-level three-dimensional graph convolution processing on the first-level three-dimensional matrix based on the first-level three-dimensional graph convolution network.
14. The method according to claim 8, characterized in that, The extracting a first-level feature sequence from the sample data based on the first network includes: Obtaining third data from the sample data; where the third data represents data of at least two dimensions corresponding to the sample data; Based on the first network, perform feature extraction on each dimension of the third data to obtain fourth data; Based on the first network, process the fourth data to obtain the first-level feature sequence.
15. An identification device, characterized in that, The recognition device includes: a first acquisition module, a first processing module, and a first recognition module; wherein: The first acquisition module is configured to acquire data to be recognized including at least two pose sequences; wherein, each pose sequence includes a plurality of node information of any pose of the target object; The first processing module is configured to perform feature extraction on the data to be recognized to obtain first-level graph convolutional data; perform first-level three-dimensional graph convolutional processing on at least two feature sequences in the first-level graph convolutional data to obtain a first-level three-dimensional convolutional result; perform first-level two-dimensional graph convolutional processing on the first-level graph convolutional data to obtain a first-level two-dimensional convolutional result; wherein, the first-level graph convolutional data includes the feature sequences of each pose sequence; the first-level three-dimensional convolutional result includes the local feature information in the first-level graph convolutional data; the first-level two-dimensional convolutional result includes the correlation relationship between the local feature information in the first-level graph convolutional data; The first recognition module is configured to recognize the pose type of the target object based on the first-level graph convolutional data, the first-level three-dimensional convolutional result, and the first-level two-dimensional convolutional result.
16. A neural network training device, characterized in that, The neural network includes a first network and a second network; wherein, the first network is used for feature extraction; the second network is used for performing graph convolutional processing on the output data of the first network; the second network at least includes a first-level three-dimensional graph convolutional network for implementing first-level three-dimensional graph convolutional processing, and a first-level two-dimensional graph convolutional network for implementing first-level two-dimensional graph convolutional processing; the device includes a second acquisition module, a second processing module, and a second recognition module; wherein: The second acquisition module is configured to acquire sample data including a plurality of pose sequences; wherein, each pose sequence includes a plurality of node information of any pose of any object; The second processing module is configured to perform feature extraction on the sample data based on the first network to obtain a first-level feature sequence; perform the first-level three-dimensional graph convolutional processing on at least two feature sequences in the first-level feature sequence based on the first-level three-dimensional graph convolutional network to obtain a first-level three-dimensional result; perform the first-level two-dimensional graph convolutional processing on the first-level feature sequence based on the first-level two-dimensional graph convolutional network to obtain a first-level two-dimensional result; wherein, the first-level feature sequence includes the feature sequences of each pose sequence; the first-level three-dimensional result includes the local feature information in the first-level feature sequence; the first-level two-dimensional result includes the correlation relationship between the local feature information in the first-level feature sequence; The second recognition module is configured to recognize the pose type of any object based on the first-level feature sequence, the first-level three-dimensional result, and the first-level two-dimensional result. The second processing module is further configured to train the first network and the second network based on the posture type to obtain a trained neural network.
17. An electronic device, characterized in that, The electronic device includes a processor, a memory, and a communication bus. The communication bus is used to establish a communication connection between the processor and the memory. The processor is configured to execute a computer program stored in the memory to implement the recognition method according to any one of claims 1-7 or the neural network training method according to any one of claims 8-14.
18. A computer-readable storage medium, characterized in that, The readable storage medium can be executed by a processor to implement the recognition method according to any one of claims 1-7 or the neural network training method according to any one of claims 8-14.