A method for skeleton behavior recognition and related devices

By extracting and aggregating the coordinate information of the key points of the skeleton, generating two-dimensional skeleton images, and using network classification models to identify features, the problem of low skeleton behavior recognition accuracy in the prior art is solved, and more efficient identification and faster operation speed is achieved.

CN115147924BActive Publication Date: 2025-06-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210762709.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2025-06-27
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

In the prior art, it is difficult to effectively extract behavioral characteristics in the time dimension of the skeleton key points in the skeleton behavior recognition, resulting in low recognition accuracy.

Method used

By obtaining the two-dimensional skeleton image, the target skeleton features are determined based on the coordinate information of each skeleton key point, and feature aggregation processing is performed to obtain the skeleton features of the target local skeleton. Then, these features are identified and processed using the preset network classification model to obtain the target classification results.

Benefits of technology

This method reduces the loss of skeleton feature information in the spatial dimension and time dimension, improves recognition accuracy, significantly reduces the amount of computing and improves the running speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115147924B_ABST
    Figure CN115147924B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method and related device for skeleton behavior recognition, which can reduce information loss of skeleton features in the spatial dimension or in the temporal dimension, improve recognition accuracy, and greatly reduce the amount of computation and improve the running speed. It at least involves technologies such as artificial intelligence. The method includes: obtaining a two-dimensional skeleton image including at least two skeleton key points; determining the target skeleton feature of each skeleton key point based on the coordinate information of each skeleton key point, and the coordinate information of the skeleton key point is used to reflect the position of the skeleton key point; performing feature aggregation processing on the target skeleton features of at least two skeleton key points to obtain the skeleton feature of the target local skeleton; performing recognition processing on the skeleton feature of the target local skeleton based on a preset network classification model to obtain a target classification result, and the target classification result is used to indicate the behavior of the target object corresponding to the skeleton.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the technical field of behavior recognition, and in particular to a method for skeleton behavior and related devices. Background Art

[0002] Based on deep learning for recognizing object behaviors in videos, it can be simply divided into two categories of "skeleton-based" and "video-based" according to the human key points of the object. Recognizing behaviors based on the skeleton key points of "skeleton-based" means constructing the skeleton information of the object by extracting the relevant information of the skeleton key points of the object, and defining the behavior activities of the people in the video through the skeleton information.

[0003] In the related solutions, the recognition of skeleton behaviors is mainly based on a convolutional neural network (CNN), that is, the skeleton key point data is spliced into two-dimensional (2D) pseudo-image data, and then the convolutional layer is used to process the spatial features of the skeleton key points to recognize the corresponding skeleton behaviors. However, this method cannot well extract the behavior features of the skeleton key points in the time dimension, resulting in a low recognition accuracy of subsequent skeleton behaviors.

[0004] Therefore, there is an urgent need for a new method for recognizing skeleton behaviors to solve the problem of poor current recognition accuracy. Summary of the Invention

[0005] The embodiments of the present application provide a method for recognizing skeleton behaviors and related devices, which can reduce the information loss of the skeleton features in the spatial dimension or the time dimension, improve the recognition accuracy, and greatly reduce the amount of computation and improve the running speed.

[0006] In a first aspect, the embodiments of the present application provide a method for recognizing skeleton behaviors. The method includes: obtaining a two-dimensional skeleton image, the two-dimensional skeleton image including at least two skeleton key points; determining the target skeleton feature of each skeleton key point based on the coordinate information of each skeleton key point, the coordinate information of the skeleton key point being used to reflect the position of the skeleton key point; performing feature aggregation processing on the target skeleton features of at least two skeleton key points to obtain the skeleton feature of the target local skeleton, the target local skeleton being composed of the skeleton key points during the feature aggregation processing; performing recognition processing on the skeleton feature of the target local skeleton based on a preset network classification model to obtain a target classification result, the target classification result being used to indicate the behavior of the target object corresponding to the skeleton.

[0007] Second aspect, an embodiment of the present application provides a skeletal behavior recognition device. The skeletal behavior recognition device includes, but is not limited to, a terminal device, a server, etc. The skeletal behavior recognition device includes an acquisition unit and a processing unit. Among them, the acquisition unit is used to acquire a two-dimensional skeletal image, and the two-dimensional skeletal image includes at least two skeletal key points. The processing unit is used to determine the target skeletal feature of each skeletal key point according to the coordinate information of each skeletal key point, and the coordinate information of the skeletal key point is used to reflect the position of the skeletal key point; perform feature aggregation processing on the target skeletal features of the skeletal key points to obtain the skeletal feature of the target local skeleton, and the target local skeleton is composed of the skeletal key points during the feature aggregation processing; perform recognition processing on the skeletal feature of the target local skeleton based on a preset network classification model to obtain a target classification result, and the target classification result is used to indicate the behavior of the target object corresponding to the skeleton.

[0008] In some optional examples, the processing unit is used to: perform feature aggregation processing on the target skeletal features of every two skeletal key points to obtain the skeletal feature of the first local skeleton, where the spatial distance between every two skeletal key points is less than or equal to a preset distance; perform feature aggregation processing on the first skeletal features of every two first local skeletons to obtain the skeletal feature of the second local skeleton, where the spatial distance between every two first local skeletons is less than or equal to a preset distance; determine the skeletal feature of the target local skeleton based on the skeletal feature of the first local skeleton and the skeletal feature of the second local skeleton.

[0009] In some other optional examples, the processing unit is used to: perform feature aggregation processing on the target skeletal feature of the first skeletal key point and the target skeletal feature of the second skeletal key point to obtain the skeletal feature of the first local skeleton, where the first skeletal key point and the second skeletal key point are any two of at least two skeletal key points whose spatial distance is less than or equal to a preset distance.

[0010] In some other optional examples, the processing unit is used to: perform feature combination processing on the target skeletal feature of the first skeletal key point and the target skeletal feature of the second skeletal key point to obtain a combined feature; perform convolution processing on the combined feature respectively based on convolutional layers of at least two convolutional scales to obtain corresponding convolutional features, where each convolutional scale is different; adjust the feature scale of each convolutional feature based on a pooling layer to obtain each adjusted convolutional feature, where the feature scale of each adjusted convolutional feature is the same; perform feature splicing processing on each adjusted convolutional feature to obtain the skeletal feature of the first local skeleton.

[0011] In some other alternative examples, the processing unit is configured to: adjust the feature dimension of the skeleton feature of the second local skeleton based on the feature dimension of the skeleton feature of the first local skeleton to obtain the skeleton feature of the adjusted second local skeleton, where the feature dimension of the skeleton feature of the adjusted second local skeleton is the same as that of the skeleton feature of the first local skeleton; perform feature aggregation processing on the skeleton feature of the first local skeleton and the skeleton feature of the adjusted second local skeleton to obtain the skeleton feature of the target local skeleton.

[0012] In some other alternative examples, the processing unit is configured to: perform deconvolution processing on the skeleton feature of the second local skeleton to obtain the deconvolved skeleton feature of the second local skeleton, where the feature scale of the deconvolved skeleton feature of the second local skeleton is the same as that of the skeleton feature of the first local skeleton; adjust the feature dimension of the deconvolved skeleton feature of the second local skeleton based on a convolutional layer to obtain the skeleton feature of the adjusted second local skeleton.

[0013] In some other alternative examples, the processing unit is configured to: divide the coordinate information into at least two coordinate slices, where each coordinate slice is used to represent the frame position of the skeleton key point at different times; determine the first position feature of the skeleton key point based on the coordinate information, where the first position feature is used to reflect the absolute position of the skeleton key point; determine the relative position information of each coordinate slice based on the coordinate difference between the coordinate information and the central coordinate information, where the central coordinate information is used to indicate the central position of the skeleton; process the relative position information of the coordinate slices based on a preset transformer model to obtain the second position feature of the skeleton key point, where the second position feature is used to reflect the relative position of the skeleton key point; perform feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton feature of each skeleton key point.

[0014] In some other alternative examples, the processing unit is further configured to, before performing feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton feature of each skeleton key point, perform feature aggregation processing on the first position feature and the second position feature to obtain the position aggregation feature of each skeleton key point; determine a relationship matrix based on the position aggregation feature of the target skeleton key point and the position aggregation features of the remaining skeleton key points, where the relationship matrix is used to indicate the similarity between the target skeleton key point and the remaining skeleton key points, and the target skeleton key point is any one of at least two skeleton key points. The acquisition unit is used to acquire a weight matrix, where the weight matrix is used to indicate the correlation degree between the coordinate slice of the target skeleton key point and the coordinate slices of the remaining skeleton key points. The processing unit is used to determine the relationship feature of the target skeleton key point according to the relationship matrix and the weight matrix to obtain the target skeleton feature of the target skeleton key point.

[0015] In some other alternative examples, the processing unit is configured to: process the position aggregation feature of the target skeleton key points based on a fully connected layer to obtain the multi-head attention information of the coordinate slices of the target skeleton key points, and process the position aggregation feature of the remaining skeleton key points based on a fully connected layer to obtain the multi-head attention information of the coordinate slices of the remaining skeleton key points; determine a relationship matrix based on the multi-head attention information of the coordinate slices of the target skeleton key points and the multi-head attention information of the coordinate slices of the remaining skeleton key points.

[0016] In some other alternative examples, the processing unit is configured to: splice the coordinate information of each skeleton key point to generate a two-dimensional skeleton image.

[0017] A third aspect of the embodiments of the present application provides a skeleton behavior recognition device, including: a memory, an input / output (I / O) interface, and a memory. The memory is used to store program instructions. The processor is configured to execute the program instructions in the memory to perform the skeleton behavior recognition method corresponding to the implementation manner of the first aspect above.

[0018] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium, in which instructions are stored, and when they run on a computer, the computer is caused to execute the method corresponding to the implementation manner of the first aspect above.

[0019] A fifth aspect of the embodiments of the present application provides a computer program product containing instructions, and when it runs on a computer or a processor, the computer or the processor is caused to execute the method corresponding to the implementation manner of the first aspect above.

[0020] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0021] In the embodiments of the present application, a two-dimensional skeleton image is obtained. The two-dimensional skeleton image includes at least two skeleton key points. Then, based on the coordinate information of each skeleton key point, the target skeleton feature of each skeleton key point is determined, and the coordinate information of the skeleton key point can reflect the position of the skeleton key point. In this way, by performing feature aggregation processing on the target skeleton features of at least two skeleton key points, the skeleton feature of the target local skeleton is obtained. The target local skeleton is composed of the skeleton key points during the feature aggregation processing. And based on a preset network classification model, the skeleton feature of the target local skeleton is recognized to obtain a target classification result. The target classification result is used to indicate the behavior of the target object corresponding to the skeleton. Through the above method, the target skeleton features of at least two skeleton key points are aggregated into the skeleton feature of the corresponding target local skeleton, thereby compressing the feature scale of the skeleton feature in the time dimension in the network, greatly reducing the amount of computation, and improving the running speed. Moreover, since the coordinate information of the skeleton key point reflects the position of the corresponding skeleton key point, then determining the target skeleton feature of each skeleton key point according to the coordinate information of each skeleton key point can realize extracting the target skeleton feature of each skeleton key point from the spatial dimension and the time dimension of the skeleton key point, thereby reducing the information loss of the skeleton feature in the spatial dimension or the skeleton feature in the time dimension, and improving the recognition accuracy accordingly. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0023] Figure 1 FIG. shows a schematic diagram of the system framework for recognizing skeleton behavior provided by the embodiments of the present application;

[0024] Figure 2 FIG. shows a flowchart of a method for skeleton behavior recognition provided by the embodiments of the present application;

[0025] Figure 3 FIG. shows a schematic diagram of a two-dimensional skeleton image provided by the embodiments of the present application;

[0026] Figure 4 FIG. shows a schematic diagram of dividing the coordinate information of skeleton key points into coordinate slices provided by the embodiments of the present application;

[0027] Figure 5 FIG. shows another flowchart of the method for skeleton behavior recognition provided by the embodiments of the present application;

[0028] Figure 6 Shows a schematic flow diagram of feature aggregation provided by an embodiment of the present application;

[0029] Figure 7 Shows a schematic structural diagram of a skeleton behavior recognition device provided by an embodiment of the present application;

[0030] Figure 8 Shows a schematic hardware structure diagram of a skeleton behavior recognition device provided by an embodiment of the present application. Detailed implementation manners

[0031] An embodiment of the present application provides a method for skeleton behavior recognition and related devices, which can reduce information loss of skeleton features in the spatial dimension or skeleton features in the time dimension, improve the recognition accuracy, and greatly reduce the amount of computation and improve the running speed.

[0032] It can be understood that in the specific implementation manner of the present application, data related to user information, etc. is involved. When the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.

[0033] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0034] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above accompanying drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0035] With the research and progress of artificial intelligence (AI) technology, AI technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, AI technology will be applied in more fields and play an increasingly important role.

[0036] Based on deep learning to identify the behavior of objects in a video, it is possible to identify the behavior of skeleton key points from the "skeleton-based" perspective, or in other words, construct the skeleton information of an object by extracting the relevant information of the skeleton key points of the object, and then identify the behavior activities performed by the object in the video through the skeleton information.

[0037] The embodiments of this application provide a method for skeleton behavior recognition. The method for skeleton behavior recognition provided by the embodiments of this application is implemented based on artificial intelligence. Artificial intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machine to have the functions of perception, reasoning, and decision-making.

[0038] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision (CV) technology, speech technology, natural language processing technology, and machine learning / deep learning.

[0039] In the embodiments of this application, the artificial intelligence technologies mainly involved include the above-mentioned directions such as machine learning and computer vision technology. For example, it can involve deep learning in machine learning (ML), including attention learning, multiayer perceptron (MLP), etc.; it can also involve video behavior recognition in computer vision technology.

[0040] The method for skeleton behavior recognition provided in this application can be applied to skeleton behavior devices with data processing capabilities, such as terminal devices, servers, etc. Among them, terminal devices can include, but are not limited to, smartphones, desktop computers, laptop computers, tablet computers, smart speakers, in-vehicle devices, smart watches, wearable smart devices, intelligent voice interaction devices, smart home appliances, aircraft, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services, etc. This application does not make specific limitations. In addition, the mentioned terminal devices and servers can be directly or indirectly connected through wired communication or wireless communication, etc. This application does not make specific limitations.

[0041] The above-mentioned skeleton behavior device can have the ability to implement the above-mentioned computer vision technology. The mentioned computer vision technology: Computer vision is a science that studies how to enable machines to "see". Further, it refers to using cameras and computers to replace the human eye to perform machine vision such as target recognition, trajectory tracing, and measurement on targets, and further perform graphics processing to make the computer process into images that are more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies and attempts to establish an artificial intelligence system that can obtain information from images or multi-dimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition. In the embodiments of this application, the skeleton behavior device can recognize the skeleton behavior of the target object through this computer vision technology.

[0042] In addition, the skeleton behavior device can also have machine learning capabilities. Machine learning is a multi-disciplinary cross-discipline that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as neural networks.

[0043] In the method for skeleton behavior recognition provided by the embodiments of the present application, an artificial intelligence model is adopted, which mainly involves the application of a neural network. The neural network is used to identify and process the skeleton features of the skeleton to obtain corresponding classification results, and then the behavior of the target object corresponding to the skeleton is indicated by the classification results.

[0044] In the traditional solution, the recognition of skeleton behavior is usually implemented based on CNN. However, CNN only considers the spatial features of the skeleton key points and cannot consider the skeleton behavior from the perspective of the behavior features in the time dimension, resulting in a low recognition accuracy of subsequent skeleton behavior.

[0045] Based on this, to solve the above-mentioned technical problems, the embodiments of the present application provide a method for skeleton behavior recognition. The method for skeleton behavior recognition can extract the target skeleton features of each skeleton key point from the spatial dimension and the time dimension of the skeleton key points, so as to reduce the information loss of the skeleton features in the spatial dimension or the skeleton features in the time dimension, thereby improving the recognition accuracy. The above-mentioned method for skeleton behavior recognition can be applied to Figure 1 the system architecture for recognizing skeleton behavior shown in. It should be understood that in actual applications, it may also be applied to other system frameworks, and the present application does not make specific limitations. Schematically, the method for skeleton behavior recognition provided by the embodiments of the present application will be introduced below in combination with the system framework.

[0046] As Figure 1 shown, for the skeleton of a certain target object, each skeleton key point in the skeleton ( Figure 1The frame information at different moments (indicated by the middle dot) is used to construct corresponding coordinate information, and then the coordinate information of all skeleton key points is used to construct a coordinate sequence matrix. Furthermore, this coordinate sequence matrix can be regarded as a two-dimensional skeleton image. In this two-dimensional skeleton image, the coordinate values in each row of the coordinate sequence can reflect the positions of the same skeleton key point in the frames at different moments. Then, for the constructed two-dimensional skeleton image, the target skeleton features corresponding to each skeleton key point can be extracted based on the coordinate information of each skeleton key point. In this way, through the feature aggregation module, the target skeleton features of at least two skeleton key points are subjected to feature aggregation processing to obtain the skeleton features of the target local skeleton. Schematically, the target skeleton features of at least two skeleton key points can be subjected to feature aggregation processing through the feature aggregation module to obtain the skeleton features of the first local skeleton, and the skeleton key points in the described first local skeleton are composed of at least two skeleton key points during the feature aggregation processing. Then, through the feature aggregation module, the skeleton features of every two first local skeletons are subjected to feature aggregation processing to obtain the skeleton features of the second local skeleton, and the skeleton key points in the described second local skeleton are composed of the skeleton key points in the two first local skeletons during the feature aggregation processing. Further, the feature aggregation module aggregates the skeleton features of the first local skeleton and the skeleton features of the second local skeleton to obtain the skeleton features of the above-mentioned target local skeleton. Finally, the skeleton features of the target local skeleton are input into a preset network classification model, and the preset network classification model performs recognition processing on the skeleton features of the target local skeleton to obtain a target classification result. In this way, by aggregating the target skeleton features of at least two skeleton key points into the skeleton features of the corresponding target local skeleton, the feature scale of the skeleton features in the network in the time dimension is gradually compressed, greatly reducing the amount of computation and improving the running speed. Moreover, the coordinate information of the skeleton key points reflects the position of the corresponding skeleton key point. Then, determining the target skeleton features of each skeleton key point according to the coordinate information of each skeleton key point can extract the target skeleton features of each skeleton key point from the spatial dimension and the time dimension of the skeleton key point, thereby reducing the information loss of the skeleton features in the spatial dimension or the time dimension, and improving the recognition accuracy.

[0047] It should be understood that in practical applications, the above Figure 1The described skeletal key points may include, but are not limited to: skeletal key points and / or joint key points, etc. The described skeletal key points may include, but are not limited to: one or more of the right shoulder key point, right elbow key point, left shoulder key point, left elbow key point, right hip key point, left hip key point, top of head key point, neck key point, right knee key point, left knee key point, right ankle key point, left ankle key point. In practical applications, it may also include left palm key point, right palm key point, etc., which are not specifically limited in the embodiments of this application. Joint key points may also include, but are not limited to: shoulder joint key point, elbow joint key point, wrist joint key point, finger joint key point, knee joint key point, ankle joint key point, etc., which are not specifically limited in the embodiments of this application.

[0048] In addition, the mentioned preset network classification model may include, but is not limited to: multi-layer perceptron (MLP) network, etc., which is not specifically limited in the embodiments of this application. The mentioned feature aggregation module can be understood as a functional module in the aforementioned skeletal behavior recognition device, and can be used to aggregate the skeletal features of at least two spatially adjacent skeletal key points into the skeletal features of an overall local skeleton.

[0049] The following introduces a method for skeletal behavior recognition provided by the embodiments of this application with reference to the accompanying drawings. Figure 2 Fig. shows a flowchart of a method for skeletal behavior recognition provided by the embodiments of this application. As Figure 2 shown, the method for skeletal behavior recognition may include the following steps:

[0050] 201. Obtain a two-dimensional skeletal image, where the two-dimensional skeletal image includes at least two skeletal key points.

[0051] In this example, for each skeletal key point, the corresponding coordinate information can be obtained. In this way, by splicing the coordinate information of each skeletal key point, a two-dimensional skeletal image is generated. For example, Figure 3 Fig. shows a schematic diagram of a two-dimensional skeletal image provided by the embodiments of this application. As Figure 3 shown, the skeletal structure diagram of the target object shows the skeletal key points at different spatial positions, such as the right elbow key point, right palm key point, left shoulder key point, right shoulder key point, neck key point, etc. If the coordinate information of each skeletal key point can be regarded as a coordinate sequence, at this time, the coordinate information of all skeletal key points can be spliced to construct a two-dimensional skeletal image, such as: Among them, each row of the coordinate sequence can be regarded as the coordinate information of a skeletal key point. For example, taking the right elbow key point, right palm key point, right shoulder key point, and neck key point as an example, {x 11 ,x 12 ,x 13 ,x14 ,..., x 1n} is regarded as the coordinate information of the right shoulder key point, and the coordinate information of the neck key point can be {x 21 , x 22 , x 23 , x 24 ,..., x 2n}, and the coordinate information of the right elbow key point is {x i1 , x i2 , x i3 , x i4 ,..., x in}, and the coordinate information of the right palm key point is {x j1 , x j2 , x j3 , x j4 ,..., x jn}, etc. In the embodiments of the present application, no specific limitation is made.

[0052] It should be noted that the two-dimensional skeleton image can be understood as a 2D pseudo-image with a shape that satisfies the parameters c, v, and n, where c represents the coordinate dimension of each skeleton key point, v represents the number of skeleton key points, and n is the number of frames of the coordinate sequence of each skeleton key point. In addition, the described skeleton key points may include, but are not limited to, one or more of bone key points and joint key points. Specifically, reference may be made to the foregoing Figure 1 for understanding, and details are not described herein.

[0053] 202. Determine the target skeleton feature of each skeleton key point based on the coordinate information of each skeleton key point. The coordinate information of the skeleton key point is used to reflect the position of the skeleton key point.

[0054] In this example, after obtaining the two-dimensional skeleton image, the shallow skeleton features of each skeleton key point in the two-dimensional skeleton image can be extracted. That is, since the coordinate information of the skeleton key point can reflect the position of the skeleton key point, the target skeleton feature of each skeleton key point can be determined through the coordinate information of each skeleton key point.

[0055] Exemplarily, for the coordinate information of each skeleton key point, the corresponding coordinate information can be divided into at least two coordinate patches, and each coordinate patch can be used to represent the frame position of the skeleton key point at different times. Each coordinate patch comes from the same skeleton key point and has similar motion characteristics. For example, Figure 4 shows a schematic diagram of dividing the coordinate information of the skeleton key point into coordinate patches provided by the embodiments of the present application. As Figure 4 shown, in the foregoing Figure 3Based on the shown two-dimensional skeleton image, the coordinate information of each skeleton key point can be divided. For example, the coordinate information {x i1 ,x i2 ,x i3 ,x i4 ,...,x in} of the right elbow key point can be divided into at least two coordinate slices, such as: {x i1 ,x i2},{x i3 ,x i4 ,},{...,x in}, etc., and the present application does not make specific limitations. As can be seen from Figure 4 , the coordinate slices in the same row all come from the same skeleton key point and can be used to reflect the frame position of the skeleton key point at different times.

[0056] Therefore, based on the coordinate information of each skeleton key point to determine the target skeleton feature of each skeleton key point, the following method can be adopted, that is: determining the first position feature of the skeleton key point based on the coordinate information, and the first position feature is used to reflect the absolute position situation of the skeleton key point; determining the relative position information of each coordinate slice based on the coordinate difference between the coordinate information and the central coordinate information, and the central coordinate information is used to indicate the central position of the skeleton. Then, processing the relative position information of the coordinate slices based on a preset transformer model to obtain the second position feature of the skeleton key point, and the second position feature is used to reflect the relative position situation of the skeleton key point. Finally, performing feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton feature of each skeleton key point.

[0057] Specifically, the coordinate information reflects the position situation of the corresponding skeleton key point. Then, based on the coordinate information of each skeleton key point, the absolute position of the skeleton key point in the spatial arrangement can be determined, that is, the first position feature of the skeleton key point can be determined. In addition, the relative position of the skeleton key point can be further considered, and this relative position can be understood as the position offset of the coordinate slices at different positions on the skeleton key point to the center of the entire skeleton. Exemplarily, taking the right elbow key point shown in Figure 4 as an example, its coordinate information {x i1 ,x i2 ,x i3 ,x i4 ,...,x in} can represent the frame position of the right elbow key point at different times, and the position of the right elbow key point in the r-th frame can be regarded as a three-dimensional vector in the spatial coordinate system, that is, {x ir ,y ir ,z ir}, 1 ≤ r ≤ n, where r and n are integers. From this coordinate information, the overall central coordinate information {μ x , μ y , μ z} can be calculated, that is: In this way, for the r-th coordinate slice in this right elbow key point, the relative position information of this coordinate slice can be calculated, that is Similarly, for each coordinate slice obtained by dividing this right elbow key point, the corresponding relative position information is determined. Then, based on a preset transformer model, feature extraction is performed on the relative position information of each coordinate slice to obtain the relative position of this right elbow key point, that is, the second position feature of the right elbow key point is obtained. In this way, the second skeleton feature of this right elbow key point is spliced along the channel dimension onto the first position feature of this right elbow key point, thereby obtaining the target skeleton feature of this right elbow key point. Through the above method, the prior information of the spatial arrangement of the right elbow key point can be introduced into the feature extraction process, improving the accuracy of feature extraction.

[0058] Similarly, for the target skeleton features of other skeleton key points, such as the extraction process of the target skeleton feature of the right palm key point, it can also be understood by referring to the extraction process of the target skeleton feature of the right elbow key point above, and will not be elaborated here.

[0059] In some examples, during the process of determining the target skeleton feature of the skeleton key point, in addition to considering the relative position and absolute position of the skeleton key point, the similarity between different skeleton key points can also be considered. If the similarity degree is higher, it indicates that the corresponding skeleton key points have similar motion behaviors to a great extent. Therefore, after dividing the coordinate information into at least two coordinate slices, the self-attention mechanism can also be used to process between different coordinate slices. Exemplarily, before aggregating the first position feature and the second position feature to obtain the target skeleton feature of each skeleton key point, the method for recognizing skeleton behavior further includes: aggregating the first position feature and the second position feature to obtain the position aggregation feature of each skeleton key point; determining a relationship matrix based on the position aggregation feature of the target skeleton key point and the position aggregation features of the remaining skeleton key points, where the relationship matrix is used to indicate the similarity between the target skeleton key point and the remaining skeleton key points, and the target skeleton key point is any one of at least two skeleton key points; obtaining a weight matrix, where the weight matrix is used to indicate the correlation degree between the coordinate slices of the target skeleton key point and the coordinate slices of the remaining skeleton key points. Then, based on the relationship matrix and the weight matrix, the relationship feature of the target skeleton key point is determined to obtain the target skeleton feature of the target skeleton key point.

[0060] In this example, after calculating the first position feature and the second position feature of the skeleton key point, the first position feature and the second position feature can be aggregated to obtain the position aggregation feature of the corresponding skeleton key point. Then, a fully connected layer is used to process the position aggregation feature of the target skeleton key point, and the multi-head attention information of the coordinate slice of the target skeleton key point is calculated. Similarly, for the remaining skeleton key points, the multi-head attention information of the coordinate slices of the corresponding skeleton key points is also calculated using fully connected layers respectively.

[0061] For example, taking the right elbow key point as the target skeleton key point, the first position feature and the second position feature of the right elbow key point can be concatenated to obtain the position aggregation feature of the right elbow key point. Then, the fully connected layer is used to process the position aggregation feature of the right elbow key point, and further obtain the multi-head attention information of each coordinate slice in the right elbow key point. Similarly, similar operations are also performed on the remaining skeleton key points such as the right palm key point, the right shoulder key point, and the neck key point, etc., to obtain the multi-head attention information of each coordinate slice in each skeleton key point. The described multi-head attention information includes the Q, K, V matrices in the multi-head attention mechanism, where, Q = F q (x), K = F k (x), V = F v (x), and x is the position aggregation feature. Then, based on the multi-head attention information of the coordinate slices of the right elbow key point and the multi-head attention information of the coordinate slices of the remaining skeleton key points, the relationship matrix is determined. In other words, after obtaining the multi-head attention information of the coordinate slices in each skeleton key point, the multi-head attention information of the coordinate slices on the same frame in different skeleton key points can be processed to obtain the relationship matrix. It should be noted that each value in the relationship matrix reflects the similarity between the target skeleton key point and the remaining skeleton key points.

[0062] To better utilize the prior relationship of the spatial arrangement between skeleton key points, a weight matrix A can be initialized. The weight matrix A is used to indicate the correlation degree between the coordinate slices of the target skeleton key point and the coordinate slices of the remaining skeleton key points, that is, each value in the weight matrix can be understood as the correlation degree between the i-th coordinate slice in the target skeleton key point and the j-th coordinate slice in another skeleton key point. It should be noted that the dimension of the weight matrix is R×W, where R is the number of coordinate slices in the target skeleton key point, and W is the number of coordinate slices in another skeleton key point.

[0063] In this way, based on the relationship matrix and the weight matrix, the relationship feature attention of the target skeleton key point can be determined, that is: where, d kBoth and T are preset values. In this way, this relational feature is used as the target skeleton feature of the target skeleton key point.

[0064] It should be noted that for the extraction process of the target skeleton features of other skeleton key points, it can also be understood with reference to the extraction process of the target skeleton features of this target skeleton key point, which will not be elaborated here.

[0065] In addition, the above-mentioned weight matrix A is a learnable parameter matrix. When initialized, each element in the weight matrix A represents the distance between the coordinate slices of the corresponding two skeleton key points. For example, taking the target skeleton key point as the right elbow key point, the element in the i-th row and j-th column of the weight matrix A is expressed as: the distance d between the i-th coordinate slice in the right elbow key point and the j-th coordinate slice in another skeleton key point (such as the right palm key point). When the i-th coordinate slice and the j-th coordinate slice are synchronized in time, this distance d can be understood as the reciprocal of the spatial distance and the reciprocal of the time distance + 1 between the right elbow key point corresponding to the i-th coordinate slice and the right palm key point corresponding to the j-th coordinate slice. By initializing the weight matrix in the above way, prior information of the human body structure can be introduced into the network, making the network converge faster and the extracted skeleton features more robust.

[0066] 203. Perform feature aggregation processing on the target skeleton features of at least two skeleton key points to obtain the skeleton features of the target local skeleton, and the target local skeleton is composed of the skeleton key points during the feature aggregation processing.

[0067] In this example, after obtaining the target skeleton features of each skeleton key point, feature aggregation processing can be performed on the target skeleton features of at least two skeleton key points to obtain the skeleton features of the target local skeleton. It should be understood that the target local skeleton is composed of the skeleton key points during the feature aggregation processing. For example, Figure 3 Taking the two-dimensional skeleton image shown as an example, since the right elbow key point and the right palm key point are adjacent in space, the target skeleton features of the right elbow key point and the right palm key point can be aggregated, and then the skeleton features of the target local skeleton can be obtained. The target local skeleton described here is composed of the right elbow key point and the right palm key point.

[0068] Specifically, Figure 5 shows another process schematic diagram of the skeleton behavior recognition method provided by the embodiment of the present application. In some optional examples, for step 203 above, it can also be specifically understood with reference to Figure 2 the content of steps S5031 to S5034 shown in Figure 5 as follows:

[0069] S5031. Aggregate the target skeleton features of every two skeleton key points to obtain the skeleton feature of the first local skeleton, where the spatial distance between every two skeleton key points is less than or equal to a preset distance.

[0070] In this example, the spatial distance between every two skeleton key points is less than or equal to the preset distance, which can be understood as the positions of every two skeleton key points with a spatial distance less than or equal to the preset distance being adjacent in space. For example, the right elbow key point and the right palm key point are adjacent in space; another example is that the left elbow key point and the left palm key point are adjacent in space, etc. The present application does not make specific limitations. In some examples, the target skeleton feature of the first skeleton key point and the target skeleton feature of the second skeleton key point can be aggregated through the Figure 1 feature aggregation module shown in to obtain the skeleton feature of the first local skeleton. The described first skeleton key point and the second skeleton key point are adjacent in space, that is, the spatial distance between the first skeleton key point and the second skeleton key point is less than or equal to the preset distance. In addition, the first local skeleton can also be understood as a mini - part - level skeleton, that is, a small part of the overall skeleton, which can be composed of the skeleton key points during the corresponding feature aggregation process. For example, it can be composed of the right elbow key point and the right palm key point.

[0071] In addition, the described feature aggregation module is composed of a group of convolutional layers (conv) with different convolutional scales and a corresponding group of pooling layers (pool). Then, in the process of aggregating the target skeleton feature of the first skeleton key point and the target skeleton feature of the second skeleton key point to obtain the skeleton feature of the first local skeleton, the target skeleton feature of the first skeleton key point and the target skeleton feature of the second skeleton key point can be first combined to obtain a combined feature. Then, the combined feature is respectively convolved based on convolutional layers with at least two convolutional scales to obtain corresponding convolutional features. It should be understood that each of these at least two convolutional scales is different. Moreover, since each convolutional scale is different, after convolution through convolutional layers with different convolutional scales, the feature scales of the corresponding convolutional features are also different. In order to aggregate convolutional features with different feature scales together, the feature scales of each convolutional feature can also be adjusted through a group of pooling layers in this feature aggregation module to obtain each adjusted convolutional feature, so that the feature scales of each adjusted convolutional feature are the same. It should be understood that each pooling layer in the group of pooling layers is connected to a convolutional layer. In this way, after adjusting to obtain convolutional features with the same feature scale for each, the adjusted convolutional features for each can be concatenated to further obtain the skeleton feature of this first local skeleton.

[0072] For example, Figure 6The flowchart of feature aggregation provided by the embodiments of the present application is shown. As Figure 6 shown, taking the right elbow key point as the first skeleton key point and the right palm key point as the second skeleton key point as an example, since the dimension of the target skeleton feature of each skeleton key point is (b, c, 1, n). Then, the feature aggregation module first combines the target skeleton features of the right elbow key point and the right palm key point that are adjacent in space in terms of the spatial sorting of the skeleton key points to form a combined feature with a dimension of (b, c, 2, n). Then, the combined feature is respectively convolved through three convolutional layers with different convolutional scales (such as: (2, 1), (2, 3), (2, 5)) and a stride of 2, so as to obtain convolutional features with different corresponding feature scales. Then, for the convolutional features with different feature scales, they are respectively pooled through the pooling layers connected to the convolutional layers to adjust the feature scales of the convolutional features. Finally, the convolutional features with the same feature scale are spliced along the channel dimension into a complete feature, that is, the skeleton feature of the first local skeleton is obtained.

[0073] It should be understood that the above Figure 5 only takes the right elbow key point as the first skeleton key point and the right palm key point as the second skeleton key point as an example to illustrate the process of skeleton feature aggregation. In practical applications, the feature aggregation module can also, based on the same aggregation process, perform skeleton feature aggregation processing on the remaining skeleton key points that are adjacent in space to obtain the skeleton features of the corresponding first local skeleton. For example, the target skeleton features of the left palm key point and the left elbow key point can be aggregated to obtain the skeleton features of the corresponding first local skeleton; or the target skeleton features of the right shoulder skeleton key point and the neck skeleton key point can be aggregated to obtain the skeleton features of the corresponding first local skeleton. The present application does not make specific limitations. In addition, the three described convolutional scales are respectively (2, 1), (2, 3), (2, 5), which are only schematic descriptions in the embodiments of the present application. In practical applications, other types of convolutional scales can also be used, and the embodiments of the present application do not make specific limitations.

[0074] S5032. Aggregate the skeleton features of every two first local skeletons to obtain the skeleton features of the second local skeleton, where the spatial distance between the skeleton key points in every two first local skeletons is less than or equal to a preset distance.

[0075] In this example, after obtaining the skeleton features of all the first local skeletons through the operation of the above step S5031, it is also possible to use the above Figure 6The described feature aggregation module performs feature aggregation processing on the skeleton features of every two first local skeletons to obtain the skeleton features of the second local skeleton. The specific process of the feature aggregation processing can be understood with reference to the content described in the above step S5031 and will not be elaborated here. It should be noted that the spatial distance between the skeleton key points in every two first local skeletons described is less than or equal to a preset distance.

[0076] It should be understood that the spatial distance between the skeleton key points in every two first local skeletons described being less than or equal to the preset distance can be understood as the positions of the two corresponding first local skeletons formed by the skeleton key points with a spatial distance less than or equal to the preset distance being adjacent in space. For example, if the right elbow key point and the right palm key point adjacent in space form the first local skeleton A, and the right shoulder skeleton key point and the neck skeleton key point adjacent in space form the first local skeleton B, then assuming the spatial distance between the first local skeleton A and the first local skeleton B is less than or equal to the preset distance, it can be considered that the first local skeleton A and the first local skeleton B are adjacent in space. In addition, the second local skeleton can also be understood as a part-level skeleton, that is, a partial skeleton in the overall skeleton, and can be composed of the skeleton key points in the first local skeleton during the corresponding feature aggregation processing.

[0077] S5033. Adjust the feature dimension of the skeleton feature of the second local skeleton based on the feature dimension of the skeleton feature of the first local skeleton to obtain the adjusted skeleton feature of the second local skeleton.

[0078] In this example, compared with the feature dimension of the skeleton feature of the first local skeleton, the feature dimension of the skeleton feature of the second local skeleton obtained after the feature aggregation processing is compressed in the time dimension. If the skeleton feature of the second local skeleton is directly input into the subsequent preset network classification model, although the network convergence speed will be very fast, a large amount of feature information will be lost in the skeleton feature of the second local skeleton, resulting in a significant reduction in the recognition accuracy during the subsequent recognition process. Therefore, through the feature upsampling module, the feature dimension of the skeleton feature of the second local skeleton can be adjusted according to the feature dimension of the skeleton feature of the first local skeleton, so that the adjusted skeleton feature of the second local skeleton is the same as the feature dimension of the skeleton feature of the first local skeleton.

[0079] Exemplarily, in the process of adjusting the feature dimension of the skeleton feature of the second local skeleton based on the feature dimension of the skeleton feature of the first local skeleton, specifically, the skeleton feature of the second local skeleton can be first deconvolved through a feature upsampling module to obtain the deconvolved skeleton feature of the second local skeleton. It should be noted that the feature scale of the described deconvolved skeleton feature of the second local skeleton is the same as the feature scale of the skeleton feature of the first local skeleton. Then, based on the convolutional layer, the feature dimension of the deconvolved skeleton feature of the second local skeleton is adjusted, so that the adjusted skeleton feature of the second local skeleton has the same feature dimension, or has the same channel dimension. It should be noted that the described feature scale can be understood as the corresponding resolution.

[0080] S5034. Perform feature aggregation processing on the skeleton feature of the first local skeleton and the adjusted skeleton feature of the second local skeleton to obtain the skeleton feature of the target local skeleton.

[0081] In this example, after obtaining the adjusted skeleton feature of the second local skeleton through the above-mentioned step S5033, the adjusted skeleton feature of the second local skeleton can be subjected to feature aggregation processing with the skeleton feature of the first local skeleton. For example, by adding the adjusted skeleton feature of the second local skeleton and the skeleton feature of the first local skeleton, the final skeleton feature of the target local skeleton can be obtained.

[0082] 204. Perform recognition processing on the skeleton feature of the target local skeleton based on a preset network classification model to obtain a target classification result, and the target classification result is used to indicate the behavior of the target object corresponding to the skeleton.

[0083] In this example, after obtaining the skeleton feature of the target local skeleton, the skeleton feature of the target local skeleton can be used as the input of the preset network classification model, so as to identify and classify the skeleton feature of the target local skeleton through the preset network model, and then obtain the corresponding target classification result.

[0084] It should be noted that the described preset network classification model may include, but is not limited to, classification models such as MLP, and is not specifically limited in the embodiments of the present application.

[0085] In the embodiments of the present application, a two-dimensional skeleton image is obtained. The two-dimensional skeleton image includes at least two skeleton key points. Then, based on the coordinate information of each skeleton key point, the target skeleton feature of each skeleton key point is determined, and the coordinate information of the skeleton key point can reflect the position of the skeleton key point. In this way, through feature aggregation processing of the target skeleton features of at least two skeleton key points, the skeleton feature of the target local skeleton is obtained. The target local skeleton is composed of the skeleton key points during the feature aggregation processing, and based on a preset network classification model, the skeleton feature of the target local skeleton is recognized to obtain a target classification result. The target classification result is used to indicate the behavior of the target object corresponding to the skeleton. Through the above method, the target skeleton features of at least two skeleton key points are aggregated into the skeleton feature of the corresponding target local skeleton, thereby compressing the feature scale of the skeleton feature in the time dimension in the network, greatly reducing the amount of computation, and improving the running speed. Moreover, since the coordinate information of the skeleton key point reflects the position of the corresponding skeleton key point, then determining the target skeleton feature of each skeleton key point according to the coordinate information of each skeleton key point can extract the target skeleton feature of each skeleton key point from the spatial dimension and the time dimension of the skeleton key point, thereby reducing the information loss of the skeleton feature in the spatial dimension or the skeleton feature in the time dimension, and improving the recognition accuracy accordingly.

[0086] The above mainly introduces the solution provided by the embodiments of the present application from the perspective of the method. It can be understood that in order to implement the above functions, it includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the modules and algorithm steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0087] The embodiments of the present application can divide the device into functional modules according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiments of the present application is illustrative, only a logical function division, and there may be other division methods in actual implementation.

[0088] The following will describe in detail the skeleton behavior recognition device in the embodiments of the present application. Figure 7This is a schematic diagram of an embodiment of the skeleton behavior recognition device provided in the embodiments of the present application. As Figure 7 shown, the skeleton behavior recognition device may include an acquisition unit 701 and a processing unit 702.

[0089] Among them, the acquisition unit 701 is used to acquire a two-dimensional skeleton image, and the two-dimensional skeleton image includes at least two skeleton key points. The processing unit 702 is used to determine the target skeleton feature of each skeleton key point according to the coordinate information of each skeleton key point, and the coordinate information of the skeleton key point is used to reflect the position of the skeleton key point; perform feature aggregation processing on the target skeleton features of the skeleton key points to obtain the skeleton feature of the target local skeleton, and the target local skeleton is composed of the skeleton key points during the feature aggregation processing; perform recognition processing on the skeleton feature of the target local skeleton based on a preset network classification model to obtain a target classification result, and the target classification result is used to indicate the behavior of the target object corresponding to the skeleton.

[0090] In some optional examples, the processing unit 702 is used to: perform feature aggregation processing on the target skeleton features of every two skeleton key points to obtain the skeleton feature of the first local skeleton, where the spatial distance between every two skeleton key points is less than or equal to a preset distance; perform feature aggregation processing on the first skeleton features of every two first local skeletons to obtain the skeleton feature of the second local skeleton, where the spatial distance between every two first local skeletons is less than or equal to a preset distance; determine the skeleton feature of the target local skeleton based on the skeleton feature of the first local skeleton and the skeleton feature of the second local skeleton.

[0091] In some other optional examples, the processing unit 702 is used to: perform feature aggregation processing on the target skeleton feature of the first skeleton key point and the target skeleton feature of the second skeleton key point to obtain the skeleton feature of the first local skeleton, where the first skeleton key point and the second skeleton key point are any two of at least two skeleton key points with a spatial distance less than or equal to a preset distance.

[0092] In some other optional examples, the processing unit 702 is used to: perform feature combination processing on the target skeleton feature of the first skeleton key point and the target skeleton feature of the second skeleton key point to obtain a combined feature; perform convolution processing on the combined feature respectively based on convolutional layers with at least two convolutional scales, where each convolutional scale is different; adjust the feature scale of each convolutional feature based on a pooling layer to obtain each adjusted convolutional feature, where the feature scales of each adjusted convolutional feature are the same; perform feature splicing processing on each adjusted convolutional feature to obtain the skeleton feature of the first local skeleton.

[0093] In some other alternative examples, the processing unit 702 is configured to: adjust the feature dimension of the skeleton feature of the second local skeleton based on the feature dimension of the skeleton feature of the first local skeleton, so as to obtain the skeleton feature of the adjusted second local skeleton, and the feature dimension of the skeleton feature of the adjusted second local skeleton is the same as that of the skeleton feature of the first local skeleton; perform feature aggregation processing on the skeleton feature of the first local skeleton and the skeleton feature of the adjusted second local skeleton to obtain the skeleton feature of the target local skeleton.

[0094] In some other alternative examples, the processing unit 702 is configured to: perform deconvolution processing on the skeleton feature of the second local skeleton to obtain the deconvolved skeleton feature of the second local skeleton, and the feature scale of the deconvolved skeleton feature of the second local skeleton is the same as that of the skeleton feature of the first local skeleton; adjust the feature dimension of the deconvolved skeleton feature of the second local skeleton based on the convolutional layer to obtain the skeleton feature of the adjusted second local skeleton.

[0095] In some other alternative examples, the processing unit 702 is configured to: divide the coordinate information into at least two coordinate slices, and each coordinate slice is used to represent the frame position of the skeleton key point at different moments; determine the first position feature of the skeleton key point based on the coordinate information, and the first position feature is used to reflect the absolute position situation of the skeleton key point; determine the relative position information of each coordinate slice based on the coordinate difference between the coordinate information and the central coordinate information, and the central coordinate information is used to indicate the central position of the skeleton; process the relative position information of the coordinate slice based on a preset transformer model to obtain the second position feature of the skeleton key point, and the second position feature is used to reflect the relative position situation of the skeleton key point; perform feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton feature of each skeleton key point.

[0096] In some other alternative examples, the processing unit 702 is further configured to, before performing feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton feature of each skeleton key point, perform feature aggregation processing on the first position feature and the second position feature to obtain the position aggregation feature of each skeleton key point; determine a relationship matrix based on the position aggregation feature of the target skeleton key point and the position aggregation features of the remaining skeleton key points, and the relationship matrix is used to indicate the similarity between the target skeleton key point and the remaining skeleton key points, and the target skeleton key point is any one of at least two skeleton key points. The obtaining unit 701 is configured to obtain a weight matrix, and the weight matrix is used to indicate the correlation degree between the coordinate slice of the target skeleton key point and the coordinate slices of the remaining skeleton key points. The processing unit 702 is configured to determine the relationship feature of the target skeleton key point according to the relationship matrix and the weight matrix to obtain the target skeleton feature of the target skeleton key point.

[0097] In some other alternative examples, the processing unit 702 is configured to: process the position aggregation feature of the target skeleton key points based on the fully connected layer to obtain the multi-head attention information of the coordinate slices of the target skeleton key points, and process the position aggregation feature of the remaining skeleton key points based on the fully connected layer to obtain the multi-head attention information of the coordinate slices of the remaining skeleton key points; determine the relationship matrix based on the multi-head attention information of the coordinate slices of the target skeleton key points and the multi-head attention information of the coordinate slices of the remaining skeleton key points.

[0098] In some other alternative examples, the processing unit 702 is configured to: splice the coordinate information of each skeleton key point to generate a two-dimensional skeleton image.

[0099] The above describes the skeleton behavior recognition device in the embodiments of the present application from the perspective of modular functional entities. The following describes the skeleton behavior recognition device in the embodiments of the present application from the perspective of hardware processing. Figure 8 It is a schematic structural diagram of the skeleton behavior recognition device provided by the embodiments of the present application. The skeleton behavior recognition device may vary greatly due to different configurations or performances. The skeleton behavior recognition device may include at least one processor 801, a communication line 807, a memory 803, and at least one communication interface 804.

[0100] The processor 801 may be a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the solution of the present application.

[0101] The communication line 807 may include a path for transmitting information between the above components.

[0102] The communication interface 804 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0103] The memory 803 can be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions. The memory can exist independently and be connected to the processor through the communication line 807. The memory can also be integrated with the processor.

[0104] Among them, the memory 803 is used to store the computer execution instructions for implementing the solution of this application, and is controlled by the processor 801 for execution. The processor 801 is used to execute the computer execution instructions stored in the memory 803, so as to implement the method for skeleton behavior recognition provided in the above embodiments of this application.

[0105] Optionally, the computer execution instructions in the embodiments of this application can also be referred to as application code, and the embodiments of this application do not make specific limitations on this.

[0106] In a specific implementation, as an embodiment, the skeleton behavior recognition device can include multiple processors, such as Figure 8 the processor 801 and the processor 802 in. Each of these processors can be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0107] In a specific implementation, as an embodiment, the skeleton behavior recognition device can further include an output device 805 and an input device 806. The output device 805 communicates with the processor 801 and can display information in various ways. The input device 806 communicates with the processor 801 and can receive inputs from the target object in various ways. For example, the input device 806 can be a mouse, a touch screen device, a sensing device, etc.

[0108] The above-mentioned skeleton behavior recognition device can be a general device or a special device. In a specific implementation, the skeleton behavior recognition device can be a server, a terminal, etc. or a device with a Figure 8 similar structure in. The embodiments of this application do not limit the type of the skeleton behavior recognition device.

[0109] It should be noted that Figure 8 the processor 801 in can make the skeleton behavior recognition device execute the method in the corresponding method embodiment by calling the computer execution instructions stored in the memory 803. Figures 2 to 6 corresponding method embodiment.

[0110] Specifically, Figure 7The function / implementation process of the processing unit 702 in can be achieved by Figure 8 the processor 801 in calling the computer-executable instructions stored in the memory 803. Figure 7 The function / implementation process of the acquisition unit 701 in can be achieved by Figure 8 the communication interface 804 in.

[0111] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0112] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0113] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0114] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0116] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0117] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product.

[0118] A computer program product includes one or more computer instructions. When the computer execution instructions are loaded and executed on a computer, they generate in whole or in part the processes or functions according to the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as an SSD), etc.

[0119] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application.

Claims

1. A method for skeletal action recognition, characterized in that, Including: Obtain a two-dimensional skeleton image, where the two-dimensional skeleton image includes at least two skeleton key points; Based on the coordinate information of each of the skeleton key points, determine the target skeleton feature of each of the skeleton key points, and the coordinate information of the skeleton key point is used to reflect the position situation of the skeleton key point; Perform feature aggregation processing on the target skeleton features of at least two of the skeleton key points to obtain the skeleton feature of the target local skeleton, where the target local skeleton is composed of the skeleton key points during the feature aggregation processing; Based on a preset network classification model, perform recognition processing on the skeleton feature of the target local skeleton to obtain a target classification result, and the target classification result is used to indicate the behavior of the target object corresponding to the skeleton; The determining the target skeleton feature of each of the skeleton key points based on the coordinate information of each of the skeleton key points includes: Divide the coordinate information into at least two coordinate slices, and each coordinate slice is used to represent the frame position of the skeleton key point at different times; Based on the coordinate information, determine the first position feature of the skeleton key point, and the first position feature is used to reflect the absolute position situation of the skeleton key point; Based on the coordinate difference between the coordinate information and the central coordinate information, determine the relative position information of each coordinate slice, and the central coordinate information is used to indicate the central position of the skeleton; Based on a preset transformer model, process the relative position information of the coordinate slice to obtain the second position feature of the skeleton key point, and the second position feature is used to reflect the relative position situation of the skeleton key point; Perform feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton feature of each of the skeleton key points.

2. The method according to claim 1, characterized in that The performing feature aggregation processing on the target skeleton features of at least two of the skeleton key points to obtain the skeleton feature of the target local skeleton includes: Perform feature aggregation processing on the target skeleton features of every two of the skeleton key points to obtain the skeleton feature of the first local skeleton, where the spatial distance between every two of the skeleton key points is less than or equal to a preset distance; Perform feature aggregation processing on the skeleton features of every two of the first local skeletons to obtain the skeleton feature of the second local skeleton, where the spatial distance between the skeleton key points in every two of the first local skeletons is less than or equal to a preset distance; Based on the skeleton feature of the first local skeleton and the skeleton feature of the second local skeleton, determine the skeleton feature of the target local skeleton.

3. The method according to claim 2, wherein The performing feature aggregation processing on the target skeleton features of every two of the skeleton key points to obtain the skeleton feature of the first local skeleton includes: Perform feature aggregation processing on the target skeleton feature of the first skeleton key point and the target skeleton feature of the second skeleton key point to obtain the skeleton feature of the first local skeleton, where the first skeleton key point and the second skeleton key point are any two of the at least two skeleton key points whose spatial distance is less than or equal to the preset distance.

4. The method according to claim 3, wherein Performing feature aggregation processing on the target skeleton features of the first skeleton key point and the target skeleton features of the second skeleton key point to obtain the skeleton features of the first local skeleton, including: Performing feature combination processing on the target skeleton features of the first skeleton key point and the target skeleton features of the second skeleton key point to obtain a combined feature; Performing convolution processing on the combined feature respectively based on convolutional layers of at least two convolutional scales to obtain corresponding convolutional features, where each of the convolutional scales is different; Adjusting the feature scales of each of the convolutional features based on a pooling layer to obtain each adjusted convolutional feature, where the feature scales of each of the adjusted convolutional features are the same; Performing feature concatenation processing on each of the adjusted convolutional features to obtain the skeleton features of the first local skeleton.

5. The method according to any one of claims 2 to 4, characterized in that, Determining the skeleton features of the target local skeleton based on the skeleton features of the first local skeleton and the skeleton features of the second local skeleton, including: Adjusting the feature dimension of the skeleton features of the second local skeleton based on the feature dimension of the skeleton features of the first local skeleton to obtain the adjusted skeleton features of the second local skeleton, where the feature dimension of the adjusted skeleton features of the second local skeleton is the same as the feature dimension of the skeleton features of the first local skeleton; Performing feature aggregation processing on the skeleton features of the first local skeleton and the adjusted skeleton features of the second local skeleton to obtain the skeleton features of the target local skeleton.

6. The method according to claim 5, wherein Adjusting the feature dimension of the skeleton features of the second local skeleton based on the feature dimension of the skeleton features of the first local skeleton to obtain the adjusted skeleton features of the second local skeleton, including: Performing deconvolution processing on the skeleton features of the second local skeleton to obtain the deconvolved skeleton features of the second local skeleton, where the feature scale of the deconvolved skeleton features of the second local skeleton is the same as the feature scale of the skeleton features of the first local skeleton; Adjusting the feature dimension of the deconvolved skeleton features of the second local skeleton based on a convolutional layer to obtain the adjusted skeleton features of the second local skeleton.

7. The method according to claim 1, characterized in that Before performing feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton features of each of the skeleton key points, the method further includes: Performing feature aggregation processing on the first position feature and the second position feature to obtain the position aggregation features of each of the skeleton key points; Determining a relationship matrix based on the position aggregation features of the target skeleton key point and the position aggregation features of the remaining skeleton key points, where the relationship matrix is used to indicate the similarity between the target skeleton key point and the remaining skeleton key points, and the target skeleton key point is any one of the at least two skeleton key points; Obtaining a weight matrix, where the weight matrix is used to indicate the correlation degree between the coordinate slices of the target skeleton key point and the coordinate slices of the remaining skeleton key points; Performing feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton features of each of the skeleton key points, including: Based on the relationship matrix and the weight matrix, determine the relationship features of the target skeleton key points to obtain the target skeleton features of the target skeleton key points.

8. The method according to claim 7, wherein The determining of the relationship matrix based on the position aggregation features of the target skeleton key points and the position aggregation features of the remaining skeleton key points includes: Process the position aggregation features of the target skeleton key points based on a fully connected layer to obtain the multi-head attention information of the coordinate slices of the target skeleton key points, and process the position aggregation features of the remaining skeleton key points based on the fully connected layer to obtain the multi-head attention information of the coordinate slices of the remaining skeleton key points; Determine the relationship matrix based on the multi-head attention information of the coordinate slices of the target skeleton key points and the multi-head attention information of the coordinate slices of the remaining skeleton key points.

9. The method according to any one of claims 1 to 4, characterized in that, The obtaining of the two-dimensional skeleton image includes: Perform splicing processing on the coordinate information of each skeleton key point to generate the two-dimensional skeleton image.

10. A skeletal behavior recognition device, characterized in that, Includes: An acquisition unit for acquiring a two-dimensional skeleton image, the two-dimensional skeleton image including at least two skeleton key points; A processing unit for determining the target skeleton features of each skeleton key point according to the coordinate information of each skeleton key point, the coordinate information of the skeleton key point being used to reflect the position of the skeleton key point; The processing unit is used to perform feature aggregation processing on the target skeleton features of the skeleton key points to obtain the skeleton features of the target local skeleton, the target local skeleton being composed of the skeleton key points during the feature aggregation processing; The processing unit is used to perform recognition processing on the skeleton features of the target local skeleton based on a preset network classification model to obtain a target classification result, the target classification result being used to indicate the behavior of the target object corresponding to the skeleton; The processing unit is specifically used for: Divide the coordinate information into at least two coordinate slices, each coordinate slice being used to represent the frame position of the skeleton key point at different times; Determine the first position feature of the skeleton key point based on the coordinate information, the first position feature being used to reflect the absolute position of the skeleton key point; Determine the relative position information of each coordinate slice based on the coordinate difference between the coordinate information and the central coordinate information, the central coordinate information being used to indicate the central position of the skeleton; Process the relative position information of the coordinate slices based on a preset transformer model to obtain the second position feature of the skeleton key point, the second position feature being used to reflect the relative position of the skeleton key point; Perform feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton features of each skeleton key point.

11. The skeleton behavior recognition device according to claim 10, wherein The processing unit is used for: Perform feature aggregation processing on the target skeleton features of every two skeleton key points to obtain the skeleton features of the first local skeleton, wherein the spatial distance between every two skeleton key points is less than or equal to a preset distance; Aggregate the first skeleton features of every two first local skeletons to obtain the skeleton features of the second local skeleton, where the spatial distance between every two first local skeletons is less than or equal to a preset distance; Determine the skeleton features of the target local skeleton based on the skeleton features of the first local skeleton and the skeleton features of the second local skeleton.

12. The skeleton behavior recognition device according to claim 11, wherein The processing unit is configured to: Aggregate the target skeleton features of the first skeleton key point and the target skeleton features of the second skeleton key point to obtain the skeleton features of the first local skeleton, where the first skeleton key point and the second skeleton key point are any two skeleton key points among the at least two skeleton key points whose spatial distance is less than or equal to the preset distance.

13. The skeleton behavior recognition device according to claim 12, characterized in that, The processing unit is configured to: Combine the target skeleton features of the first skeleton key point and the target skeleton features of the second skeleton key point to obtain a combined feature; Perform convolution processing on the combined feature respectively based on convolutional layers of at least two convolutional scales to obtain corresponding convolutional features, where each of the convolutional scales is different; Adjust the feature scales of each of the convolutional features based on a pooling layer to obtain each adjusted convolutional feature, where the feature scales of each of the adjusted convolutional features are the same; Perform feature concatenation processing on each of the adjusted convolutional features to obtain the skeleton features of the first local skeleton.

14. The skeleton behavior recognition device according to any one of claims 11 to 13, characterized in that, The processing unit is configured to: Adjust the feature dimension of the skeleton features of the second local skeleton based on the feature dimension of the skeleton features of the first local skeleton to obtain the adjusted skeleton features of the second local skeleton, and the feature dimension of the adjusted skeleton features of the second local skeleton is the same as the feature dimension of the skeleton features of the first local skeleton; Aggregate the skeleton features of the first local skeleton and the adjusted skeleton features of the second local skeleton to obtain the skeleton features of the target local skeleton.

15. The skeleton behavior recognition device according to claim 14, characterized in that, The processing unit is configured to: Perform deconvolution processing on the skeleton features of the second local skeleton to obtain the deconvolved skeleton features of the second local skeleton, and the feature scale of the deconvolved skeleton features of the second local skeleton is the same as the feature scale of the skeleton features of the first local skeleton; Adjust the feature dimension of the deconvolved skeleton features of the second local skeleton based on a convolutional layer to obtain the adjusted skeleton features of the second local skeleton.

16. The skeletal behavior recognition device according to claim 10, wherein The processing unit is further configured to: Before aggregating the first position feature and the second position feature to obtain the target skeleton features of each of the skeleton key points, aggregate the first position feature and the second position feature to obtain the position aggregation features of each of the skeleton key points; Determine a relationship matrix based on the position aggregation features of the target skeleton key point and the position aggregation features of the remaining skeleton key points, where the relationship matrix is used to indicate the similarity between the target skeleton key point and the remaining skeleton key points, and the target skeleton key point is any one of the at least two skeleton key points. The obtaining unit is configured to obtain a weight matrix, which is used to indicate the correlation degree between the coordinate slices of the target skeleton key points and the coordinate slices of the remaining skeleton key points; Performing feature aggregation processing on the first position feature and the second position feature to obtain the target skeleton feature of each skeleton key point, including: The processing unit is configured to determine the relationship feature of the target skeleton key point based on the relationship matrix and the weight matrix, so as to obtain the target skeleton feature of the target skeleton key point.

17. The skeleton behavior recognition device according to claim 16, characterized in that, The processing unit is configured to: Process the position aggregation feature of the target skeleton key point based on a fully connected layer to obtain the multi-head attention information of the coordinate slices of the target skeleton key point, and process the position aggregation feature of the remaining skeleton key points based on the fully connected layer to obtain the multi-head attention information of the coordinate slices of the remaining skeleton key points; Determine the relationship matrix based on the multi-head attention information of the coordinate slices of the target skeleton key point and the multi-head attention information of the coordinate slices of the remaining skeleton key points.

18. The skeletal behavior recognition device according to any one of claims 10 to 13, characterized in that The processing unit is configured to: Perform splicing processing on the coordinate information of each skeleton key point to generate the two-dimensional skeleton image.

19. A skeleton behavior recognition device, characterized in that, Including: An input / output (I / O) interface, a processor, and a memory, wherein program instructions are stored in the memory; The processor is configured to execute the program instructions stored in the memory and execute the method according to any one of claims 1 to 9.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions, and when the instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 9.

21. A computer program product, characterized in that, The computer program product includes instructions, and when the instructions run on a computer device, the computer device is caused to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Human skeleton-oriented motion prediction method and system

    CN111199216A

  • Action recognition method and device based on multi-feature fusion of key points, medium and equipment

    CN114299615A