Skeleton behavior identification method and device for multi-scale fusion token

By introducing a multi-scale fusion token in skeleton behavior recognition, multi-level embedded features are generated and reconstruction learning is carried out, the prediction inaccuracy problem caused by insufficient comprehensive feature information in the unsupervised method is solved, and the accuracy and reliability of behavior prediction are improved.

CN120014397APending Publication Date: 2025-05-16CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510112952.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing unsupervised skeleton behavior recognition methods are not comprehensive enough in the representation of characteristic information, resulting in insufficient accuracy in predicting human behavior.

Method used

A multi-scale fusion token is proposed to identify the skeleton behavior. By mapping the original skeleton sequence to the high-dimensional feature space and dividing the time dimension, multi-level embedded features are generated, including spatial fusion features, temporal fusion features and sequence fusion features. Different random masking strategies are used to mask multi-level embedded features, and reconstruction learning and fusion feature enhancement are performed through the encoder decoder structural model, and the behavioral action type is determined by using the encoder and special classifiers.

Benefits of technology

By considering multi-level information, the accuracy and reliability of prediction of human behavior are improved, and the prediction inaccuracy caused by the existing methods is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014397A_ABST
    Figure CN120014397A_ABST
Patent Text Reader

Abstract

The invention provides a skeleton behavior identification method and device for a multi-scale fusion token, and relates to the technical field of computers, and the method comprises the steps: mapping an original skeleton sequence to a high-dimensional feature space, and carrying out the time dimension division of the mapped original skeleton sequence, and obtaining a plurality of sequence tuples; adding three fusion features in each tuple and among the tuples to generate multi-level embedded features; the multi-level embedded features comprise joint-level features, space fusion features, time fusion features and sequence fusion features; performing corresponding mask operation on the multi-level embedded features by using different mask strategies, and performing reconstruction learning and fusion feature enhancement through a coder-decoder structure model; and determining the behavior action type of the to-be-identified target by using an encoder and a special classification head suitable for the fusion features. According to the method, the influence of multi-level information on the human body behaviors is considered, so that the accuracy and reliability of predicting the human body behaviors are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a skeleton behavior recognition method and device for multi-scale fusion tokens. Background Art

[0002] An important topic in computer vision is to study human movement to express purpose, status, etc. Skeleton behavior recognition aims to predict human behavior through the three-dimensional spatial coordinates of human joints. It has a wide range of applications, such as virtual reality, video surveillance, human-computer interaction, and smart medicine. The most common way to do this is to predict human movements from small video sequences. Since human behavior is diverse, it not only involves single-person behavior, but also multi-person coordinated behavior, making the task highly complex. In addition, for the same action, the action data obtained by different actors and different perspectives are different, making this task very challenging.

[0003] Traditionally, fully supervised methods are used, but this requires a lot of labor costs and is very time-consuming, so it is particularly important to use unsupervised methods. The unsupervised method uses masked joints as the input of the encoder and then reconstructs the mask through the decoder. This method ignores the importance of other levels of information for understanding human behavior, making the prediction of human behavior through reconstruction learning inaccurate. Summary of the invention

[0004] The purpose of the present invention is to solve the problem that the current unsupervised method does not fully represent the feature information, thus resulting in inaccurate prediction of human behavior. The present invention provides a skeleton behavior recognition method and device with multi-scale fusion tokens.

[0005] The technical solution of the embodiment of the present application is implemented as follows:

[0006] A first aspect of an embodiment of the present application provides a skeleton behavior recognition method of a multi-scale fusion token, comprising:

[0007] Mapping the original skeleton sequence to a high-dimensional feature space, and dividing the mapped original skeleton sequence by time dimension to obtain multiple sequence tuples;

[0008] Add three fusion features within each tuple and between tuples to generate multi-level embedded features; the multi-level embedded features include spatial fusion features, temporal fusion features and sequence fusion features, and the sequence fusion features are obtained by fusing information in the time domain and the space domain;

[0009] Using different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and reconstructing learning and fusion feature enhancement on the encoder-decoder structure model;

[0010] The encoder and a special classifier suitable for fusion features are used to determine the type of behavior action of the target to be identified.

[0011] Optionally, three fusion features are added within each tuple and between tuples to generate multi-level embedded features, including:

[0012] Introduce additional spatial fusion markers to aggregate all spatial information of each frame and obtain spatial fusion features;

[0013] Introduce additional time fusion markers to aggregate all the time information of each joint and obtain the time fusion feature;

[0014] The temporal fusion feature and the spatial fusion feature are combined to generate a sequence fusion feature.

[0015] Optionally, using different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and performing reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, includes:

[0016] During masking, only joint-level features, spatial fusion features, and temporal fusion are masked, while sequence fusion features are not masked;

[0017] The multi-level embedded features are input to the spatial decoder, adding masked spatial fusion tokens;

[0018] The multi-level embedded features are input to the temporal decoder, augmented with masked temporal fusion tokens.

[0019] Optionally, using different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and performing reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, includes:

[0020] Reconstruction losses in the spatial domain and the temporal domain are obtained, and the multi-level embedding features are adjusted using the reconstruction losses.

[0021] Optionally, using different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and performing reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, includes:

[0022] A knowledge distillation loss for the spatial domain and the temporal domain is obtained, and the multi-level embedding features are adjusted using the distillation loss.

[0023] Optionally, using different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and performing reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, includes:

[0024] The temporal fusion feature and the spatial fusion feature are fused to generate a sequence fusion feature, and a contrast loss of the fusion process is obtained, and the multi-level embedding feature is adjusted using the contrast loss.

[0025] Optionally, the step of using an encoder and a special classifier suitable for fusion features to determine the behavior action type of the target to be identified further includes:

[0026] The encoder and a special classifier suitable for fusion features are used to determine the behavior action type of the target to be identified.

[0027] The second aspect of the embodiment of the present application provides a skeleton behavior recognition device of a multi-scale fusion token, comprising: a sequence mapping module, a fusion generation module, a feature learning module and a behavior determination module, wherein:

[0028] The sequence mapping module is configured to map the original skeleton sequence to a high-dimensional feature space, and divide the mapped original skeleton sequence into a time dimension to obtain a plurality of sequence tuples;

[0029] The fusion generation module is configured to use different random masking strategies within each tuple and between tuples to perform corresponding masking operations on pre-randomly initialized spatial fusion tokens and temporal fusion tokens to generate multi-level embedded features; the multi-level embedded features include spatial fusion features, temporal fusion features and sequence fusion features, and the sequence fusion features are obtained by fusion information of the time domain and the space domain;

[0030] The feature learning module is configured to use different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and to perform reconstruction learning and fusion feature enhancement on the encoder-decoder structure model;

[0031] The behavior determination module is configured to determine the behavior action type of the target to be identified by using an encoder and a special classifier suitable for fusion features.

[0032] A third aspect of an embodiment of the present application provides an electronic device, comprising a processor and a memory; the memory stores a computer program, wherein the computer program, when executed by the processor, implements the skeleton behavior recognition method of the multi-scale fusion token described in the first aspect.

[0033] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0034] Compared with the prior art, the technical solution provided by this application has the following beneficial effects:

[0035] The present invention provides a skeleton behavior recognition method and device of multi-scale fusion tokens, which maps the original skeleton sequence to a high-dimensional feature space, and divides the mapped original skeleton sequence into time dimensions to obtain multiple sequence tuples; adds three fusion tokens within each tuple and between tuples to generate multi-level embedded features; uses different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and performs reconstruction learning and fusion feature enhancement through an encoder-decoder structure model; uses an encoder and a special classifier suitable for fusion features to determine the behavior action type of the target to be identified. By considering the impact of multi-level information on human behavior, the accuracy and reliability of predicting human behavior are improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 A schematic flow chart of a skeleton behavior recognition method of a multi-scale fusion token provided in an embodiment of the present application;

[0037] Figure 2 A schematic diagram of the structure of a feature converter composed of basic blocks provided in an embodiment of the present application;

[0038] Figure 3 A schematic diagram of a joint adoption scenario provided by an embodiment of the present invention;

[0039] Figure 4 A schematic diagram of the structure of a skeleton behavior recognition device for a multi-scale fusion token provided in an embodiment of the present application;

[0040] Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0041] Below, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present application.

[0042] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. The terms "include", "comprising", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.

[0043] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification, and should not be interpreted in an idealized or overly rigid manner.

[0044] Some block diagrams and / or flow charts are shown in the accompanying drawings. It should be understood that some blocks or combinations thereof in the block diagrams and / or flow charts may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that these instructions, when executed by the processor, may create a device for implementing the functions / operations described in these block diagrams and / or flow charts.

[0045] In some embodiments, see Figure 1 , Figure 1 A schematic flow chart of a skeleton behavior recognition method for a multi-scale fusion token provided in an embodiment of the present application; the skeleton behavior recognition method for a multi-scale fusion token provided in an embodiment of the present application includes:

[0046] S110, mapping the original skeleton sequence to a high-dimensional feature space, and dividing the mapped original skeleton sequence by time dimension to obtain a plurality of sequence tuples.

[0047] For example, the original skeleton sequence Mapping to high-dimensional feature space These embedded features are then divided into n tuples (or segments) in the time dimension. Each tuple represents the skeleton operation information within a certain period of time. N = nv, then the shape and size of the skeleton sequence after grouping is

[0048] S120, adding three fusion features within each tuple and between tuples to generate multi-level embedded features; the multi-level embedded features include joint-level features, spatial fusion features, temporal fusion features and sequence fusion features, and the sequence fusion features are obtained by fusing information in the time domain and the space domain.

[0049] In some embodiments, S120, three fusion features are added within each tuple and between tuples to generate multi-level embedded features, including:

[0050] Introduce additional spatial fusion markers to aggregate all spatial information of each frame and obtain spatial fusion features;

[0051] Introduce additional time fusion markers to aggregate all the time information of each joint and obtain the time fusion feature;

[0052] The temporal fusion feature and the spatial fusion feature are combined to generate a sequence fusion feature.

[0053] In one example, an additional spatial fusion tag is introduced to aggregate the spatial information of each frame and improve the model's discrimination ability, for example, when one action is the reverse of another action in time. Similarly, the concept of temporal fusion in the temporal domain is similar. In addition, the sequence fusion tag enhances the recognition of samples that are different in time and space by combining temporal and spatial features, thereby increasing the diversity of the model representation. Specifically, the spatial fusion feature Temporal fusion features and sequence fusion features Add to S through formula (1):

[0054]

[0055] Among them, concat(*) represents the concatenation operation, G and Z are the intermediate variables generated, and X is the feature variable finally input to the encoder.

[0056] S130, using different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and performing reconstruction learning and fusion feature enhancement through an encoder-decoder structure model.

[0057] In an example, given two spatial masking methods {m j1 ,m j2} and two time mask modes {m f1 ,m f2}, thus, the three mask sequences generated and The mask mode used is m j1 ∪m f1 、m j1 ∪m f2 and m j2 ∪m f1 Where T m and V m represents the length of the data frame and the number of joints after masking, and ∪ represents the joint operator. Then, F is calculated according to the corresponding mask patterns of the three mask sequences. spa Hehe F tmpFinally, the masked fusion features are concatenated with the mask sequence to obtain the masked multi-level feature embedding. mr For example, F spa Using m f1 Mask strategy gets F tmp Using m j1 The mask strategy is obtained Using formula (1), we can get the splicing method of formula (2):

[0058]

[0059] The superscript m indicates a mask operation.

[0060] X m Reconstructed through a transformer-based encoder-decoder structure. The encoder can be composed of D e The temporal attention block is composed of a number of spatiotemporal attention blocks, and each spatiotemporal attention block is composed of a spatial attention block and a temporal attention block. Specifically, for the spatial attention block, the temporal dimension is treated as a batch through the reshaping operation, and then the dependencies between joints are learned through the basic attention network composed of a multi-head attention layer (Multi-Head Self-Attention, MSA) and a feed forward neural network (Feed Forward Network, FNN) to obtain the global features of all joint points, which are then reshaped to the original shape. The temporal attention block is similar in structure to the spatial attention block, except that the dimension of the operation during reshaping becomes the spatial dimension. It should be noted that when the mask data is first input to the spatial or temporal attention block, the corresponding learnable position embedding is added.

[0061] In some embodiments, S130, using different random masking strategies to perform corresponding masking operations on the multi-level embedded features, and performing reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, including:

[0062] Obtain reconstruction losses for the spatial and temporal domains, and use the reconstruction losses to adjust the multi-level embedding features.

[0063] In this embodiment, the decoder consists of two parts: a spatial decoder and a temporal decoder. The spatial decoder consists of spatial attention blocks, and the temporal decoder consists of temporal attention blocks, and the depth of both is set to D d Similarly, spatially learnable position embeddings are added before feature data is fed into the spatial decoder, and similarly for the temporal decoder. Let S be the original skeleton sequence, S ′ is the prediction result, Δ represents S and S ′The difference between each pair of joint points between S and S is calculated. In the calculation of reconstruction loss, the smooth L1 function is used. Therefore, S and S ′ The mathematical expression of the loss between is as follows:

[0064]

[0065] Due to the simultaneous masking in the spatial and temporal domains, two types of reconstruction losses need to be measured. First, the reconstruction loss of the spatial mask component is calculated:

[0066]

[0067] Among them, T v is the number of frames without mask.

[0068] Similarly, the reconstruction loss count formula for the time domain is as follows:

[0069]

[0070] Among them, T m is the number of masked frames. Finally, the complete reconstruction loss is as follows:

[0071]

[0072] In some embodiments, S130, using different random masking strategies to perform corresponding masking operations on the multi-level embedded features, and performing reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, including:

[0073] The knowledge distillation loss of spatial and temporal domains is obtained, and the reconstruction loss is used to adjust the multi-level embedding features.

[0074] By applying m on the sequence S j1 ∪m f1 and m j2 ∪m f1 The masking strategy generates the mask sequence S mr and S ms With the same time mask strategy m f1 Therefore, from S mr and S ms The obtained spatial fusion features should be consistent. To ensure this expectation, knowledge distillation is performed on the two spatial fusion features. The obtained spatial fusion embedding is input into a fully connected (FC) layer to further regress the fusion features. The knowledge distillation loss of the model in the spatial domain is Defined as:

[0075]

[0076] in, and is the mask sequence S mr and S ms The spatial fusion features after regression. The knowledge distillation of fusion features in the time domain is similar to that of spatial fusion features. Therefore, the total loss of knowledge distillation is:

[0077]

[0078] In some embodiments, S130, using different random masking strategies to perform corresponding masking operations on the multi-level embedded features, and performing reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, including:

[0079] The temporal fusion features and the spatial fusion features are fused to generate sequence fusion features, and the contrast loss of the fusion process is obtained. The contrast loss is used to adjust the multi-level embedding features.

[0080] When training the model, sequence-level features can be obtained by fusing information in the time domain and the spatial domain. Sequence-level information not only includes global features, but also can better adapt to the recognition of downstream tasks. In order to ensure these two prerequisites, a discrimination task (i.e., contrastive learning) is performed on the sequence-level fusion markers, and contrastive loss is obtained. Positive samples come from the output of different masking strategies for the same sequence, while negative samples are composed of different input sequences. The specific process is as follows. First, a fully connected layer is used to obtain the sequence-level embedding P for sequence fusion information. n It should be noted that P n Contains the spatiotemporal information of each person in the sequence. Generally speaking, there are two people in a sequence, so the average is taken to obtain the final sequence-level embedding x, which is expressed as follows:

[0081]

[0082] Among them, N represents the number of people contained in the sequence. Finally, the InfoNCE loss is used to calculate the loss of contrastive learning, that is:

[0083]

[0084] Among them, x i is the true value of the i-th sequence, z is the corresponding positive sample, B is the batch size, and τ is the temperature hyperparameter. It should be noted that contrastive learning in the network does not require additional data enhancement and memory banks, which can reduce data calculation and memory consumption for storing negative samples.

[0085] In addition, the same sequence can have three mask modes. The view in the reconstruction learning is regarded as the true value, while the other views are regarded as positive samples. The total loss of contrastive learning can be expressed as:

[0086]

[0087] in, Calculated S mr and S ms The sequence level loss between YesS mr and S mt The loss between.

[0088] S140, using multi-level embedded features to reconstruct and learn the encoder-decoder structure model and enhance features to determine the behavior action type of the target to be identified.

[0089] In this embodiment, the three-level feature fusion is obtained through the previously pre-trained encoder, namely, the spatial fusion feature Temporal fusion features And sequence fusion features Among them C e Indicates the number of channels of the encoder output. In order to make full use of these features for skeleton-based action recognition, a special classifier is designed for the fusion features of these three levels. It should be noted that F spa and F tmp are similar. Therefore, tmp Take this as an example to illustrate the details of the classification implementation.

[0090] See also Figure 2 , Figure 2 A structural diagram of a feature transformer composed of basic blocks provided in an embodiment of the present application; a learnable position embedding is added to F tmp In the example, we obtain the position-aware embedding F t i mp , and then input into Figure 2 In the feature transformer composed of the basic blocks shown in the figure, it should be noted that, unlike the traditional Transformer, the regularization B is introduced to force the model to learn the common patterns in the skeleton sequence. Its mathematical formula is as follows:

[0091]

[0092] Among them, Q, K, and V are respectively It is projected into three fully connected layers.

[0093] Will By D s After the attention blocks, a FC layer is used to calculate and predict the scores of each behavior. spa The processing method is similar to F tmpIt should be noted that regularization is only used to calculate F tmp The attention map of F is used to learn the intrinsic pattern of the joint points, but there is no such property in the time dimension, so it is not correct to use F spa Regularization is added. In addition, in order to improve the accuracy of the model, a simple multilayer perceptron (MLP) is applied to the fusion features at the sequence level, and the final classification score is obtained by adding the scores of the fusion features at the three levels.

[0094] In an optional embodiment, Python code is written on an Intel(R) i79700K CPU (3.6GHz, 8 cores) and 16GB RAM device, running on an Ubuntu 22.04.5 LTS64 system, and tested on the NTU series datasets. The NTU series datasets include two datasets, NTU-RGB 60 and NTU-RGB 120. Both data are collected through the Mircosoft Kinect sensor device to collect the 3D spatial coordinates of the positions of 25 joint points, specifically corresponding to the joint positions as follows Figure 3 As shown, Figure 3 Schematic diagram of the joint adoption scenario provided for an embodiment of the present invention. NTU-RGB-D 60 contains a total of 56,880 data segments, with a maximum of 2 people in each action segment. The NTU-RGB-D 120 dataset is a larger version of NTU-RGB 60, with a total of 114,480 action segments. Before the experiment, each data segment contained 300 frames of data, and the data was downsampled to 64 frames of data through bilinear interpolation. During the training process, the cumulative batch size and the maximum number of training times were set to 128 and 150 respectively, and the AdamW optimizer with a learning rate of 0.0004 was used to optimize the model. Compared with the prior art, the motion recognition results under similar test conditions and environments are more accurate and more robust.

[0095] In some embodiments, see Figure 4 , Figure 4 A schematic diagram of the structure of a skeleton behavior recognition device for a multi-scale fusion token provided in an embodiment of the present application; The present application embodiment provides a skeleton behavior recognition device 400 for a multi-scale fusion token, including: a sequence mapping module 410, a fusion feature generation module 420, a feature learning module 430 and a behavior determination module 440, wherein:

[0096] A sequence mapping module 410 is configured to map the original skeleton sequence to a high-dimensional feature space, and divide the mapped original skeleton sequence into a time dimension to obtain a plurality of sequence tuples;

[0097] The fusion generation module 420 is configured to add three fusion features within each tuple and between tuples to generate a multi-level embedded feature; the multi-level embedded feature includes a spatial fusion feature, a temporal fusion feature and a sequence fusion feature, and the sequence fusion feature is obtained by fusing information in the time domain and the space domain;

[0098] A feature learning module 430 is configured to use different random masking strategies to implement corresponding masking operations on multi-level embedded features, and to perform reconstruction learning and fusion feature enhancement through an encoder-decoder structure model;

[0099] The behavior determination module 440 is configured to determine the behavior action type of the target to be identified by using an encoder and a special classifier suitable for fusion features.

[0100] In some embodiments, the fusion generation module 420 is specifically configured as follows:

[0101] Introduce additional spatial fusion markers to aggregate all spatial information of each frame and obtain spatial fusion features;

[0102] Introduce additional time fusion markers to aggregate all the time information of each joint and obtain the time fusion feature;

[0103] The temporal fusion feature and the spatial fusion feature are combined to generate a sequence fusion feature.

[0104] In some embodiments, the feature learning module 430 is specifically configured as follows:

[0105] During masking, only joint-level features, spatial fusion features, and temporal fusion are masked, while sequence fusion features are not masked;

[0106] The multi-level embedded features are input to the spatial decoder, adding masked spatial fusion tokens;

[0107] The multi-level embedded features are input to the temporal decoder, augmented with masked temporal fusion tokens.

[0108] In some embodiments, the feature learning module 430 is specifically configured as follows:

[0109] Obtain reconstruction losses for the spatial and temporal domains, and use the reconstruction losses to adjust the multi-level embedding features.

[0110] In some embodiments, the feature learning module 430 is specifically configured as follows:

[0111] The knowledge distillation loss of spatial and temporal domains is obtained, and the distillation loss is used to adjust the multi-level embedding features.

[0112] In some embodiments, the feature learning module 430 is specifically configured as follows:

[0113] The temporal fusion features and the spatial fusion features are fused to generate sequence fusion features, and the contrast loss of the fusion process is obtained. The contrast loss is used to adjust the multi-level embedding features.

[0114] In some embodiments, the behavior determination module 440 is specifically configured to:

[0115] The encoder and a special classifier suitable for fusion features are used to determine the type of behavior action of the target to be identified.

[0116] The skeleton behavior recognition device for multi-scale fusion tokens provided in the embodiments of the present application can implement the various processes in the embodiments corresponding to the skeleton behavior recognition method for multi-scale fusion tokens mentioned above, and will not be described again here to avoid repetition.

[0117] It should be noted that the skeleton behavior recognition device of the multi-scale fusion token provided in the embodiment of the present application and the skeleton behavior recognition method of the multi-scale fusion token provided in the embodiment of the present application are based on the same application concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the skeleton behavior recognition method of the aforementioned multi-scale fusion token, and the repeated parts will not be repeated.

[0118] In some embodiments, see Figure 5 , Figure 5 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. An electronic device 500 provided in an embodiment of the present application includes a processor 510 and a memory 520; the memory 520 stores a computer program, wherein the computer program implements the above-mentioned skeleton behavior recognition method of multi-scale fusion tokens when executed by the processor.

[0119] Specifically, the processor 510 may include, for example, a general-purpose microprocessor, an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 510 may also include an onboard memory for cache purposes. The processor 510 may be a single processing unit or multiple processing units for executing different actions of the method flow according to the embodiment of the present application.

[0120] The memory 520 may be any medium capable of containing, storing, conveying, propagating or transmitting instructions. For example, the memory 520 may include, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared or semiconductor system, device, component or propagation medium. Specific examples of the memory 520 include: a magnetic storage device, such as a magnetic tape or a hard disk (HDD); an optical storage device, such as a compact disk (CD-ROM); a random access memory (RAM) or flash memory; and / or a wired / wireless communication link.

[0121] The present application also provides a computer-readable medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for recognizing skeleton behavior of multi-scale fusion tokens. The computer-readable medium may be included in the device / apparatus / system described in the above-mentioned embodiment; or it may exist independently without being assembled into the device / apparatus / system. The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed, the method according to the embodiment of the present application is implemented.

[0122] According to an embodiment of the present application, a computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that may be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, wherein a computer-readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wired, optical cable, radio frequency signal, etc., or any suitable combination of the above.

[0123] Those skilled in the art will appreciate that the features described in the various embodiments and / or claims of the present application may be combined and / or combined in a variety of ways, even if such combinations or combinations are not explicitly described in the present application. In particular, without departing from the spirit and teachings of the present application, the features described in the various embodiments and / or claims of the present application may be combined and / or combined in a variety of ways. All of these combinations and / or combinations fall within the scope of the present application. Therefore, the scope of the present application should not be limited to the above-described embodiments, but should be determined not only by the appended claims, but also by the equivalents of the appended claims.

Claims

1. A skeleton behavior recognition method based on multi-scale fusion tokens, characterized in that: include: Mapping the original skeleton sequence to a high-dimensional feature space, and dividing the mapped original skeleton sequence by time dimension to obtain multiple sequence tuples; Add three kinds of fusion tokens within each tuple and between tuples to generate multi-level embedding features; The multi-level embedding features include joint-level features, spatial fusion features, temporal fusion features and sequence fusion features, and the sequence fusion features are obtained by fusing information in the temporal domain and the spatial domain; Different random masking strategies are used to implement corresponding masking operations on multi-level embedded features, and reconstruction learning and fusion feature enhancement are performed through the encoder-decoder structure model; The encoder and a special classifier suitable for fusion features are used to determine the type of behavior action of the target to be identified.

2. The skeleton behavior recognition method of multi-scale fusion token according to claim 1 is characterized in that: The method adds three fusion tokens within each tuple and between tuples to generate multi-level embedding features, including: Introduce additional spatial fusion markers to aggregate all spatial information of each frame and obtain spatial fusion features; Introduce additional time fusion markers to aggregate all the time information of each joint and obtain the time fusion feature; The temporal fusion feature and the spatial fusion feature are combined to generate a sequence fusion feature.

3. The skeleton behavior recognition method of multi-scale fusion token according to claim 1 is characterized in that: The method uses different random mask strategies to implement corresponding mask operations on multi-level embedded features, and performs reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, including: During masking, only joint-level features, spatial fusion features, and temporal fusion are masked, while sequence fusion features are not masked; The multi-level embedded features are input to the spatial decoder, adding masked spatial fusion tokens; The multi-level embedded features are input to the temporal decoder, augmented with masked temporal fusion tokens.

4. The skeleton behavior recognition method of multi-scale fusion token according to claim 1 is characterized in that: The method uses different random mask strategies to implement corresponding mask operations on multi-level embedded features, and performs reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, including: Reconstruction losses in the spatial domain and the temporal domain are obtained, and the multi-level embedding features are adjusted using the reconstruction losses.

5. The skeleton behavior recognition method of multi-scale fusion token according to claim 1 is characterized in that: The method uses different random mask strategies to implement corresponding mask operations on multi-level embedded features, and performs reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, including: A knowledge distillation loss for the spatial domain and the temporal domain is obtained, and the multi-level embedding features are adjusted using the distillation loss.

6. The skeleton behavior recognition method of multi-scale fusion token according to claim 1 is characterized in that: The method uses different random mask strategies to implement corresponding mask operations on multi-level embedded features, and performs reconstruction learning and fusion feature enhancement through an encoder-decoder structure model, including: The temporal fusion feature and the spatial fusion feature are fused to generate a sequence fusion feature, and a contrast loss of the fusion process is obtained, and the multi-level embedding feature is adjusted using the contrast loss.

7. The skeleton behavior recognition method of multi-scale fusion token according to claim 1 is characterized in that: The method of using an encoder and a special classifier suitable for fusion features to determine the behavior action type of the target to be identified also includes: After the encoder is used to learn the three fusion features, the behavior action type of the target to be identified is determined by a special classifier.

8. A skeleton behavior recognition device with multi-scale fusion tokens, characterized in that: include: Sequence mapping module, fusion generation module, feature learning module and behavior determination module, among which, The sequence mapping module is configured to map the original skeleton sequence to a high-dimensional feature space, and divide the mapped original skeleton sequence into a time dimension to obtain a plurality of sequence tuples; The fusion generation module is configured to add three fusion features within each tuple and between tuples to generate multi-level embedded features; the multi-level embedded features include spatial fusion features, temporal fusion features and sequence fusion features, and the sequence fusion features are obtained by fusion information of the time domain and the space domain; The feature learning module is configured to use different random masking strategies to implement corresponding masking operations on the multi-level embedded features, and to perform reconstruction learning and fusion feature enhancement on the encoder-decoder structure model; The behavior determination module is configured to determine the behavior action type of the target to be identified by using an encoder and a special classifier suitable for fusion features.

9. An electronic device comprising a processor and a memory; the memory stores a computer program, wherein: When the computer program is executed by the processor, the computer program implements the skeleton behavior recognition method of the multi-scale fusion token according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.