A General Action Recognition Method, Device, Computer Equipment and Readable Storage Medium Based on Semantic Feature and Structural Feature Matching
By combining semantic and structural feature matching methods, the problem of insufficient accuracy of traditional action recognition methods is solved, and a more efficient action recognition effect is achieved.
Patent Information
- Application Number
- CN202411460328.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-10-18
AI Technical Summary
Traditional action recognition methods often rely on a single feature, resulting in insufficient recognition accuracy.
Using a method based on semantic features and structural features matching, the human body position is positioned through the human body recognition model, and combined with the action autocoding model and the human body pose estimation model, the action semantics and structural features are extracted to match the most similar action categories from the benchmark feature library.
It improves the accuracy and versatility of action recognition, comprehensively considers semantic and structural characteristics, and improves the reliability of recognition results.
Smart Images

Figure CN119274237B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology. Specifically, it relates to a general action recognition method, device, computer device and readable storage medium based on the matching of semantic features and structural features. Background Art
[0002] In today's digital age, action recognition is of great significance in many fields, such as security monitoring, intelligent human-computer interaction, sports analysis, etc. Traditional action recognition methods often have limitations. For example, they rely only on a single feature for recognition, resulting in insufficient accuracy. Summary of the Invention
[0003] The purpose of the present invention is to provide a general action recognition method, device, computer device and readable storage medium based on the matching of semantic features and structural features.
[0004] In a first aspect, an embodiment of the present invention provides a general action recognition method based on the matching of semantic features and structural features, including:
[0005] Obtain a to-be-tested picture, and input the to-be-tested picture into a pre-trained human body recognition model to obtain a human body position area;
[0006] Call a pre-trained action autoencoder model to extract features from the human body position area to obtain to-be-tested action semantic features;
[0007] Call a pre-trained human body pose estimation model to extract features from the human body position area to obtain to-be-tested action structural features;
[0008] Obtain the benchmark action semantic features and benchmark action structural features corresponding to multiple benchmark pictures from a pre-set benchmark feature library, and each of the benchmark pictures is marked with a corresponding benchmark action category;
[0009] According to the multiple benchmark action semantic features and the multiple benchmark action structural features, perform feature matching on the to-be-tested action semantic features and the to-be-tested action structural features, determine the target benchmark picture with the highest matching degree, and use the target benchmark action category corresponding to the target benchmark picture as the action recognition result of the to-be-tested picture.
[0010] In a possible implementation manner, the action autoencoder model is obtained through the following method, including:
[0011] Construct an initial action auto-encoder model, where the initial action auto-encoder model includes five cascaded downsampling components and five upsampling components. The downsampling component has a Conv-BN-PRelu-Conv-BN-PRelu structure, and the upsampling component has a DeConv-BN-PRelu-DeConv-BN-PRelu structure. Each of the downsampling components and each of the upsampling components are configured with a channel attention component, and the channel attention component includes a global pooling layer, a fully connected layer, and a normalization layer;
[0012] Train the initial action auto-encoder model based on an unsupervised mechanism until a preset mean squared error loss function training termination condition is reached to obtain the action auto-encoder model.
[0013] In a possible implementation manner, the step of calling a pre-trained human pose estimation model to extract features from the human position area to obtain the to-be-tested action structure features includes:
[0014] Call a pre-trained human pose estimation model to perform key point detection on the human position area to obtain multiple human key point position information;
[0015] According to the multiple human key point position information, determine the corresponding adjacency matrix and angle matrix;
[0016] According to the preset key point weights, the adjacency matrix, and the angle matrix, calculate the to-be-tested action structure features.
[0017] In a possible implementation manner, the step of calculating the to-be-tested action structure features according to the preset key point weights, the adjacency matrix, and the angle matrix includes:
[0018] According to the formula: Calculate the to-be-tested action structure features;
[0019] where, is the to-be-tested action structure feature, is the adjacency matrix, and the adjacency matrix is used to represent the relationship between human key points, is the angle matrix used to represent the horizontal angle of the line connecting two human key points pairwise,
[0020] In a possible implementation, the step of performing feature matching on the semantic features and structural features of the action to be measured according to the multiple semantic features of the reference actions and the multiple structural features of the reference actions, and determining the target reference picture with the highest matching degree, and using the target reference action category corresponding to the target reference picture as the action recognition result of the picture to be measured includes:
[0021] Calculate the semantic cosine similarity between the semantic features of the action to be measured and the semantic features of each reference action;
[0022] Calculate the structural cosine similarity between the structural features of the action to be measured and the structural features of each reference action;
[0023] Determine whether there is a pending reference picture for which both the semantic cosine similarity and the structural cosine similarity are greater than a preset cosine similarity threshold;
[0024] If there is, use the pending reference picture as the target reference picture, and use the target reference action category corresponding to the target reference picture as the action recognition result of the picture to be measured;
[0025] If not, determine the action recognition result of the picture to be measured as an unknown action category.
[0026] In a possible implementation, the method further includes:
[0027] In the case where there are multiple pending reference pictures, calculate the Euclidean distances between the pending semantic features and pending structural features corresponding to each pending reference picture and the semantic features and structural features of the action to be measured;
[0028] According to the calculation results of the Euclidean distances, determine the pending reference picture with the highest matching degree as the target reference picture.
[0029] In a possible implementation, the step of obtaining the picture to be measured includes:
[0030] Obtain the picture to be measured and the service scenario corresponding to the picture to be measured;
[0031] The step of obtaining the semantic features of the reference actions and the structural features of the reference actions corresponding to multiple reference pictures from a pre-set reference feature library includes;
[0032] Obtain the semantic features of the reference actions and the structural features of the reference actions corresponding to multiple reference pictures from the pre-set reference feature library according to the service scenario.
[0033] In a second aspect, an embodiment of the present invention provides a general action recognition device based on semantic feature and structural feature matching, including:
[0034] An acquisition module, configured to acquire a to-be-tested picture, input the to-be-tested picture into a pre-trained human body recognition model to obtain a human body position area; call a pre-trained action auto-encoding model to perform feature extraction on the human body position area to obtain to-be-tested action semantic features; call a pre-trained human body pose estimation model to perform feature extraction on the human body position area to obtain to-be-tested action structure features; acquire benchmark action semantic features and benchmark action structure features corresponding to a plurality of benchmark pictures from a pre-set benchmark feature library, and each of the benchmark pictures is marked with a corresponding benchmark action category;
[0035] An identification module, configured to perform feature matching on the to-be-tested action semantic features and the to-be-tested action structure features according to the plurality of benchmark action semantic features and the plurality of benchmark action structure features, determine a target benchmark picture with the highest matching degree, and use the target benchmark action category corresponding to the target benchmark picture as the action recognition result of the to-be-tested picture.
[0036] In a third aspect, an embodiment of the present invention provides a computer device, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method described in the first aspect.
[0037] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program runs, it controls a computer device where the readable storage medium is located to execute the method described in the first aspect.
[0038] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a general action recognition method, device, computer device and readable storage medium based on semantic feature and structure feature matching disclosed by the present invention, by acquiring a human body position area obtained by a human body recognition model for a to-be-tested picture, then respectively extracting to-be-tested action semantic and structure features through an action auto-encoding model and a human body pose estimation model, acquiring semantic and structure features of benchmark actions from a benchmark feature library, performing feature matching based on the two to determine a target benchmark picture, and using the corresponding action category as the recognition result of the to-be-tested picture, comprehensively improving the accuracy and generality of action recognition by combining semantic and structure features. Description of the Drawings
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 Schematic diagram of the step process of the general action recognition method based on semantic feature and structure feature matching provided by an embodiment of the present invention;
[0041] Figure 2 Schematic diagram of the framework of the general action recognition method based on semantic feature and structure feature matching provided by an embodiment of the present invention;
[0042] Figure 3 Schematic diagram of the overall structure of the auto - encoding model provided by an embodiment of the present invention;
[0043] Figure 4 Schematic diagram of the component structure of the auto - encoding model provided by an embodiment of the present invention;
[0044] Figure 5 Schematic diagram of human key points provided by an embodiment of the present invention;
[0045] Figure 6 Schematic block diagram of the structure of the general action recognition device based on semantic feature and structure feature matching provided by an embodiment of the present invention;
[0046] Figure 7 Schematic block diagram of the structure of the computer device provided by an embodiment of the present invention. Detailed implementation manners
[0047] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.
[0048] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings.
[0049] To solve the technical problems in the foregoing background art, Figure 1 Schematic diagram of the process of the general action recognition method based on semantic feature and structure feature matching provided by an embodiment of the present disclosure. The following will introduce the general action recognition method based on semantic feature and structure feature matching in detail.
[0050] Step S201: Obtain a to - be - measured picture, and input the to - be - measured picture into a pre - trained human body recognition model to obtain a human body position area;
[0051] Step S202: Invoke a pre - trained action auto - encoding model to extract features from the human body position area to obtain to - be - measured action semantic features;
[0052] Step S203: Invoke a pre-trained human pose estimation model to extract features from the human position area, obtaining the structural features of the action to be measured.
[0053] Step S204: Obtain the semantic features and structural features of the reference actions corresponding to multiple reference pictures from a pre-set reference feature library. Each of the reference pictures is labeled with the corresponding reference action category.
[0054] Step S205: According to the multiple semantic features of the reference actions and the multiple structural features of the reference actions, perform feature matching on the semantic features of the action to be measured and the structural features of the action to be measured, determine the target reference picture with the highest matching degree, and use the target reference action category corresponding to the target reference picture as the action recognition result of the picture to be measured.
[0055] In an embodiment of the present invention, by way of example, it is assumed that the server is processing a monitoring video analysis task of a large shopping mall to identify various actions of customers in the mall. The server obtains one by one the pictures to be measured from the monitoring system of the mall. These pictures to be measured contain the scenes of various areas of the mall, such as the pictures of different locations such as the mall entrance, corridors, and inside the stores.
[0056] The server has previously trained a human body recognition model, which is trained based on a large amount of image data containing human bodies. When the server inputs the obtained pictures to be measured into this human body recognition model, the model starts the recognition work.
[0057] For example, for a picture to be measured at the mall entrance, the model will detect the position areas of all people in the picture. If there is a family of three walking into the mall in the picture, the model will accurately mark the body contour ranges of these three people respectively, that is, the human position areas. This area can be represented by a rectangular box or an irregular polygon, which accurately encloses the position of the human body in the picture and provides a basis for subsequent action analysis.
[0058] The training process of the human body recognition model involves deep learning algorithms. It learns various features of the human body through technologies such as convolutional neural networks (CNNs). During the training process, a large number of labeled human body images are input into the network, and the network continuously adjusts its internal weight parameters to improve the recognition accuracy of the human body. When facing new pictures to be measured, the model can accurately locate the human position area according to the learned feature patterns.
[0059] In the mall monitoring scenario, after the server obtains the human position areas of the above-mentioned family of three, it will invoke the pre-trained action auto-encoding model to extract the semantic features of the actions.
[0060] This action auto-encoding model is constructed and trained according to a specific structure. For each of the three family members, the model focuses on the image information within their body position area.
[0061] For example, for the body position area of the father, if he is carrying a shopping bag and walking with a slight forward lean, the action auto-encoding model will deeply analyze the image pixels, textures, etc. within this area. Through its own downsampling component and upsampling component structure, it gradually extracts the semantic features related to this action. The Conv-BN-PRelu-Conv-BN-PRelu structure of the downsampling component will perform convolution operations, normalization processes, etc. on the image to reduce the dimension of the image while retaining key information. The global pooling layer in the channel attention component will comprehensively statistically analyze the features of the entire area, and the fully connected layer and normalization layer will further process these statistical information, enabling the model to pay attention to the importance differences between different channels.
[0062] Through such a series of complex operations, the action auto-encoding model extracts a unique semantic feature of the action to be measured for the father walking while carrying a shopping bag. This feature can be a high-dimensional vector, which contains various semantic representations related to the action semantics, such as the overall rhythm of the action, the smoothness of the posture, etc.
[0063] The initial action auto-encoding model constructed by the action auto-encoding model contains five cascaded downsampling components and five upsampling components. This structure helps to gradually extract and restore image features. The structure of the downsampling component can effectively reduce the resolution of the image and extract key features through operations such as convolution and activation functions. The channel attention component is to enable the model to better pay attention to the important information in different channels of the image, thereby improving the accuracy of feature extraction.
[0064] When training this action auto-encoding model, it is trained based on an unsupervised mechanism. Unsupervised training means that there is no need to manually label information such as action categories for the data. The model continuously adjusts its own parameters by minimizing the preset mean square error loss function. For example, in the training dataset, there are a large number of human body images with different actions. The model will try to reconstruct the features of the input image so that the mean square error between the reconstructed features and the original features is minimized. When this error reaches the preset training termination condition, it is considered that the model training is completed, and thus it can accurately extract the action semantic features of the new human body position area.
[0065] Continuing with the example of this family of three in the mall, the server calls the pre-trained human pose estimation model to extract features from their body position areas.
[0066] For the mother's body position area, if she is reaching for a product on the shelf. The human pose estimation model first performs key point detection, which can identify the position information of multiple key parts of the mother's body, such as the positions of the head, shoulders, arm joints, waist, knees, etc. The position information of these key parts is like the coordinate points of the human pose.
[0067] Suppose the model detects that the mother's right arm is straight and her left hand is at her side. Based on this position information of multiple human key points, the model will determine the corresponding adjacency matrix and angle matrix. The adjacency matrix can represent the connection relationship between each key point. For example, the shoulder and the arm joint are connected. The angle matrix represents the horizontal angle of the line connecting two key points. For example, when the arm is straight, the horizontal angle of the line connecting the shoulder to the arm joint.
[0068] Then, according to the preset key point weights (this weight can be determined based on experience or during the model training process. For example, the weight of the arm joint may be relatively high because arm movements are more important in judging the overall movement structure), the adjacency matrix, and the angle matrix, the structural feature of the action to be measured for the mother to reach for the product is calculated. This structural feature can be a numerical vector, which reflects the structural relationship of each part of the mother's body in this action.
[0069] The human pose estimation model uses deep learning algorithms, especially variants of convolutional neural networks (CNNs), to achieve key point detection. During the training process, the model learns the key point patterns in a large amount of human pose data. By performing convolutional operations, pooling operations, etc. on the input image of the body position area, the positions of each key point are gradually located.
[0070] The determination of the adjacency matrix and the angle matrix is based on the physiological structure and kinematic principles of the human body. There are certain connection relationships between each key point of the human body, and these relationships will show different geometric structures under different actions. By calculating the connection relationship (adjacency matrix) and angle relationship (angle matrix) between these key points, and combining the preset key point weights, a vector that can comprehensively reflect the action structure feature can be obtained. This calculation method can accurately capture the structural changes of the human body under different actions.
[0071] In the scenario of mall monitoring, a large number of benchmark pictures with pre-labeled action categories are stored in the benchmark feature library of the server. These benchmark pictures cover various typical action scenarios that may occur in the mall.
[0072] For example, there is a reference picture marked with the reference action category of "selecting goods". This picture shows a customer standing in front of a shelf, looking at the goods with their eyes, and making the action of selecting the goods with both hands. The server will obtain the reference action semantic features and reference action structure features corresponding to this picture from the reference feature library.
[0073] For the semantic features, it contains information related to the semantics of the action of selecting goods, which can be extracted through a previous similar action auto-encoding model. For the structure features, they are obtained by analyzing the human pose in the picture through a human pose estimation model. For example, features related to the action structure such as the customer's body bending slightly towards the shelf and both hands moving within a certain range.
[0074] Similarly, there are also reference pictures marked with different reference action categories such as "paying" and "consulting the store clerk". The server will obtain their corresponding reference action semantic features and reference action structure features from the reference feature library.
[0075] The reference feature library is established during the system construction and pre-training stage. When building this library, first, a large number of pictures covering various action scenarios are collected, and these pictures are accurately labeled with action categories. Then, the same action auto-encoding model and human pose estimation model as those for the pictures to be tested are used to extract features from these reference pictures, obtaining the reference action semantic features and reference action structure features respectively, and storing these features together with the corresponding reference action categories in the reference feature library. In this way, when performing action recognition on the pictures to be tested, necessary reference features can be obtained from this library for comparative analysis.
[0076] Still taking a family of three in a shopping mall as an example, assume that the server has extracted the semantic features and structural features of the action to be tested for the father walking while carrying a shopping bag, and has obtained the reference action semantic features and reference action structure features of many reference pictures from the reference feature library.
[0077] First, the server calculates the semantic cosine similarity between the semantic features of the action to be tested and each reference action semantic feature. For example, for the reference action semantic features corresponding to the reference action category of "walking normally", the semantic cosine similarity between them is calculated. At the same time, the structural cosine similarity between the semantic features of the action to be tested and each reference action structure feature is calculated.
[0078] Assume that the semantic cosine similarity and structural cosine similarity corresponding to the reference action category of "walking normally" are both relatively high and both are greater than the preset cosine similarity threshold. Then this reference picture will be tentatively designated as a pending reference picture.
[0079] If there is only this one pending reference picture, then it will be determined as the target reference picture, and the corresponding reference action category of "normal walking" will be used as the action recognition result of the father in this picture to be tested.
[0080] However, if there are multiple pending reference pictures, for example, there is also a reference picture corresponding to the reference action category of "quickly walking towards the destination" that also meets the condition that the semantic and structural cosine similarities are greater than the threshold. In this case, the server will calculate the Euclidean distances between the pending action semantic features and pending action structural features corresponding to each pending reference picture and the to-be-tested action semantic features and to-be-tested action structural features.
[0081] Suppose that after calculation, the Euclidean distance corresponding to the pending reference picture of "normal walking" is smaller, which means that its overall matching degree with the action of the father walking while carrying a shopping bag is higher. Then the pending reference picture of "normal walking" will be determined as the target reference picture, and the corresponding reference action category of "normal walking" is the final action recognition result.
[0082] If when calculating the semantic cosine similarity and structural cosine similarity, none of the reference pictures meet the condition that both are greater than the preset cosine similarity threshold, then the server will determine the action recognition result of this picture to be tested as an unknown action category. For example, if the father makes a very strange action and there is no reference action in the reference feature library with a high matching degree, it will be marked as an unknown action category.
[0083] The semantic cosine similarity is an index for measuring the similarity degree between two vectors in the semantic space. By calculating the cosine value of the included angle between the to-be-tested action semantic feature vector and the reference action semantic feature vector, the semantic similarity between them can be obtained. The structural cosine similarity is the same in principle and is used to measure the similarity degree between action structural features.
[0084] The preset cosine similarity threshold is determined based on experience or in the model training and verification stage, and it is used to screen out possible matching reference pictures. The Euclidean distance is to calculate the straight-line distance between two vectors in space. When there are multiple pending reference pictures, the smaller the Euclidean distance, the closer the two feature vectors are as a whole, so that the target reference picture with the highest matching degree can be determined more accurately. This multi-stage and multi-index feature matching method can improve the accuracy and reliability of action recognition.
[0085] In the embodiment of the present invention, the action auto-encoding model can be obtained in the following manner and can be implemented through the following examples.
[0086] Construct an initial action autoencoder model, where the initial action autoencoder model includes five cascaded downsampling components and five upsampling components. The downsampling component has a Conv-BN-PRelu-Conv-BN-PRelu structure, and the upsampling component has a DeConv-BN-PRelu-DeConv-BN-PRelu structure. Each of the downsampling components and each of the upsampling components is configured with a channel attention component, and the channel attention component includes a global pooling layer, a fully connected layer, and a normalization layer;
[0087] Train the initial action autoencoder model based on an unsupervised mechanism until a preset mean squared error loss function training termination condition is reached to obtain the action autoencoder model.
[0088] In an embodiment of the present invention, exemplarily, the server starts to construct an initial action autoencoder model. This model is specifically designed to be able to effectively extract the semantic features of human actions.
[0089] It includes five cascaded downsampling components and five upsampling components. The design of this structure is based on the requirements of image feature extraction and restoration. The downsampling component is used to gradually reduce the resolution of the input image (here is the human position area image), and at the same time extract the key features in the image; the upsampling component is used to gradually restore the features after downsampling to a feature representation close to the original image, so as to finally obtain a feature vector containing rich semantic information.
[0090] Downsampling component - Conv-BN-PRelu-Conv-BN-PRelu structure:
[0091] Convolutional layer (Conv):
[0092] For the first convolutional layer in the downsampling component, the server will set appropriate parameters such as the convolutional kernel size and stride when constructing the model. For example, the convolutional kernel size may be set to 3×3 and the stride is 1. When the human position area image is input, this convolutional layer will perform a convolution operation on each pixel point in the image and its surrounding pixels. Taking the shopping mall monitoring scenario as an example, if the input is a human position area image of a customer bending down to select goods, the convolutional layer will identify local features in the image, such as the texture of the customer's clothes and the local shape of the body contour. It slides the convolutional kernel on the image and calculates the convolution result at each position, thereby initially extracting some basic features related to the action.
[0093] The second convolutional layer further performs a convolution operation on the output of the first convolutional layer. For example, it can further refine the feature extraction, perhaps conduct a more in-depth analysis of features such as the body contour and clothing texture extracted by the first convolutional layer, such as identifying the relationship between the folds on the clothing and the body's bending movements, etc.
[0094] Batch Normalization layer (BN):
[0095] After each convolutional layer is the Batch Normalization layer. When processing image data of a large number of customers in a mall, due to factors such as the lighting and shooting angles of different images, the distribution of the input data may vary. The Batch Normalization layer normalizes the features output by the convolutional layer. For example, for images of a group of customers under different lighting conditions, the Batch Normalization layer normalizes the feature values after convolution for each batch of images, so that their means and variances are within a certain range, thereby improving the training efficiency and stability of the model.
[0096] For consecutive action images of the same customer, due to possible slight changes in lighting or minor jitter of the shooting device, etc., the Batch Normalization layer can also ensure the consistency of the model's feature processing for these images.
[0097] PRelu activation function:
[0098] After each convolutional layer and Batch Normalization layer is the PRelu activation function. The PRelu activation function introduces non-linearity to the features. In a mall scenario, when the model processes images of customers in different postures, such as a customer's movement from standing to squatting, the shape and posture changes of the body are non-linear. The PRelu activation function can better capture this non-linear relationship, enabling the model to better learn the feature representations in different action states. Without an activation function, the model can only handle linear relationships and cannot accurately represent complex human action features.
[0099] Upsampling component - DeConv - BN - PRelu - DeConv - BN - PRelu structure:
[0100] Deconvolution layer (DeConv):
[0101] The deconvolution layer in the upsampling component is the inverse operation of the convolutional layer in the downsampling component. In a mall surveillance scenario, assuming that the customer action features processed by the downsampling component are low-resolution representations, the deconvolution layer gradually increases the resolution of the feature map. For example, for a customer bending action feature that has been downsampled to a smaller size, the deconvolution layer expands this low-resolution feature according to the pattern it has learned, attempting to restore it to a representation close to the original image features.
[0102] The second deconvolution layer further refines this restoration process, gradually restoring more information in the original image, such as the overall contour of the customer's body, details of the action posture, etc.
[0103] Batch Normalization layer (BN) and PRelu activation function
[0104] Similar to the downsampling component, the Batch Normalization layer and PRelu activation function in the upsampling component also play important roles. The Batch Normalization layer ensures the stability of features during the deconvolution process, avoiding unstable model training due to changes in data distribution. The PRelu activation function further introduces non-linearity to the features after deconvolution, enabling the model to better reconstruct the complex action semantic features in the original image.
[0105] Channel attention component - Global pooling layer, Fully connected layer, and Normalization layer:
[0106] Global pooling layer:
[0107] A channel attention component is configured in each downsampling component and upsampling component. The global pooling layer is the first operation of the channel attention component. Taking the customer image in mall surveillance as an example, when an image of the customer's body position area is input and partially processed by the downsampling component, the global pooling layer performs a pooling operation on the entire feature map in the spatial dimension. It compresses the entire feature map into a vector in the channel dimension, which contains the comprehensive information of the entire feature map in each channel. For example, for a feature map with multiple channels (which may represent different feature dimensions, such as features of different body parts, different direction features of actions, etc.), the global pooling layer will statistically process all pixel values in each channel (such as taking the average) to obtain a value reflecting the overall feature of the entire feature map in that channel, and finally form a vector in the channel dimension.
[0108] Fully connected layer:
[0109] After receiving the vector in the channel dimension output by the global pooling layer, the fully connected layer performs a linear transformation on this vector. In the mall scenario, the fully connected layer will readjust the importance of different channels according to the patterns that the model has learned. For example, if a certain channel is very important for representing the action of the customer reaching for a commodity (such as the channel representing the direction of the arm movement), the fully connected layer will increase the weight of this channel in the overall feature representation; while for some less important channels (such as channels that may be related to background noise), it will reduce their weights.
[0110] Normalization layer:
[0111] The normalization layer normalizes the results output by the fully connected layer. When processing a large amount of image data of different customers and different actions in a shopping mall, the normalization layer can ensure that the features of different images after passing through the channel attention component are consistent in the numerical range. In this way, in subsequent model operations, the combination and processing of features can be carried out more stably, improving the model's ability to accurately extract the semantic features of different actions.
[0112] The server collected a large amount of unlabeled image data for training the initial action autoencoder model. The sources of these image data are extensive, including frame images containing human actions intercepted from various public image datasets, network videos, etc. In the scenario of shopping mall monitoring, it can also include some internal monitoring video images of the shopping mall with unlabeled action categories.
[0113] For example, these images contain various actions of people in different environments, such as walking in the park, walking on the street, and performing various daily activities indoors. These image data do not need to be pre-labeled with action categories because an unsupervised training mechanism is adopted.
[0114] During the training process, the goal of the initial action autoencoder model is to gradually compress the features of the input image through the downsampling component, and then try to reconstruct the feature representation of the input image as accurately as possible through the upsampling component. For each input unlabeled image (such as an image of a person bending down to pick up something in a room), the model first compresses the image features through the downsampling component. During the downsampling process, through the cooperation of the convolutional layer, batch normalization layer, and activation function, the high-resolution features of the image are gradually converted into low-resolution feature representations, while extracting key action features.
[0115] Then, the upsampling component attempts to reconstruct the features of the original image based on the features obtained by downsampling. During this process, the transposed convolutional layer, batch normalization layer, and activation function work together to gradually restore the low-resolution features to a feature representation close to the original image. During the reconstruction process, the channel attention component will continuously adjust the weights of each channel to better focus on the channels that have an important impact on the action semantic features.
[0116] The server calculates the preset mean squared error loss function. The mean squared error loss function is used to measure the difference between the features reconstructed by the model and the features of the original input image. For example, for an input image containing the action of a customer reaching for a product, there may be differences between the features reconstructed by the model and the features of the original image at each pixel point or feature dimension. The mean squared error loss function will calculate the average of the sum of the squares of these differences.
[0117] During the training process, the server continuously inputs new unlabeled images into the model, calculates the value of the mean squared error loss function, and adjusts the model's parameters based on this value. For example, if the mean squared error between the reconstructed features and the original features for a certain image is large, it indicates that the model's feature extraction and reconstruction effects for this image are not good. The model will adjust the parameters of the convolutional layers, batch normalization layers, activation functions in the downsampling and upsampling components, as well as the fully connected layer in the channel attention component through the backpropagation algorithm to reduce the mean squared error.
[0118] As the unlabeled image data is continuously input for training, the value of the mean squared error loss function will gradually decrease. When the value of this mean squared error loss function reaches the preset training termination condition, for example, when the mean squared error drops to a very small value (such as 0.001), it indicates that the model has been able to extract and reconstruct the features of the input image well. At this time, it is considered that the model training is completed, and an action autoencoder model is obtained. This action autoencoder model can be used to extract the action semantic features of the images in the customer body position area in mall monitoring.
[0119] In the embodiment of the present invention, to extract the features of the human body position area by calling the pre-trained human pose estimation model to obtain the to-be-detected action structure features, it can be implemented through the following examples.
[0120] Call the pre-trained human pose estimation model to perform key point detection on the human body position area to obtain the position information of multiple human key points;
[0121] According to the position information of the multiple human key points, determine the corresponding adjacency matrix and angle matrix;
[0122] According to the preset key point weights, the adjacency matrix, and the angle matrix, calculate to obtain the to-be-detected action structure features.
[0123] In the embodiment of the present invention, for example, assume that the server is processing a video surveillance screen of a sports event to identify the action postures of athletes. Each frame of image obtained from the video is used as the to-be-detected picture. After obtaining the human body position area (for example, the human body contour area of the athlete on the track in a track and field competition scene), the server calls the pre-trained human pose estimation model to perform key point detection.
[0124] This human pose estimation model is trained based on a large amount of image data containing human pose annotations, in which the accurate positions of multiple key points such as the head, neck, shoulders, elbows, wrists, chest, waist, knees, and ankles are annotated.
[0125] For the human body position area of track and field athletes, the model first identifies the position of the head, which can be determined by searching for specific patterns in the image, such as the outline of the hair, facial features, etc. Then, the model accurately locates the position of the shoulders, which may be determined by comparing with surrounding body parts (such as the neck and arm parts). Similarly, for the elbows and wrists, the model detects their positions based on the shape of the arm and the characteristics of the joints. The position of the waist may be determined by the curve of the body and the connection relationship with the upper and lower body parts, and the knees and ankles are located according to the bending and shape characteristics of the legs.
[0126] At the moment when a high jumper takes off, the model can accurately detect the positions of various key points of his body. The head is in an upwardly raised posture, the shoulders are open and raised upwards, the elbows are bent, the wrists are in a specific position as the arms swing, the waist is bent to prepare for the takeoff force, the knees are bent to store energy, and the ankles provide support and elasticity for the takeoff. The position information of these key points is represented in the form of coordinates. For example, in the image coordinate system, each key point has corresponding x and y coordinate values, thus accurately describing the positions of various parts of the athlete's body at the moment of takeoff.
[0127] The human pose estimation model is usually built based on a convolutional neural network (CNN). During the training process, a large number of labeled human pose images are input, and the model gradually learns the feature patterns of each key point in different human poses through operations such as convolutional layers and pooling layers.
[0128] For example, in the convolutional layer, convolutional kernels of different sizes scan the image to extract local features. These local features are downsampled through the pooling layer, reducing the amount of data while retaining key information. As the network deepens, the model can learn more abstract features, thus being able to accurately identify the positions of key points in different human poses.
[0129] During the detection process, the model inputs the human body position area image into the network. After a series of convolutional and pooling operations, the position information of each key point is output through the final fully connected layer or convolutional layer, and this information is represented in the form of coordinate values for the position in the image.
[0130] Continuing with the example of the high jumper's takeoff moment, the server determines the adjacency matrix and angle matrix based on the position information of multiple detected human key points.
[0131] For the adjacency matrix, the physiological structure of the human body determines the natural connection relationships between various key points. For example, the head is connected to the neck, the neck is connected to the shoulders, the shoulders are connected to the elbows, the elbows are connected to the wrists, the shoulders are also connected to the chest, the chest is connected to the waist, the waist is connected to the knees, and the knees are connected to the ankles. According to this connection relationship, the elements in the adjacency matrix can be represented as 0 or 1. If two key points are connected, the corresponding element is 1; otherwise, it is 0. This matrix clearly represents the connection relationships between various key points.
[0132] For the angle matrix, the server calculates the horizontal angles of the lines connecting pairs of key points. For example, for two key points, the shoulder (S) and the elbow (E), taking the shoulder as the origin, a coordinate system is established, and the angle between the line connecting the shoulder to the elbow and the horizontal direction (x-axis) is calculated. If the athlete's arm is bent upward at the moment of takeoff, this angle can be an acute angle, assumed to be 30 degrees. Similarly, the angles between other pairs of key points, such as the neck and the shoulders, the elbows and the wrists, etc., are calculated.
[0133] In a dance performance scenario, when the dancer makes complex movement postures, the server will also determine the adjacency matrix and the angle matrix of the human body key points. For different dance movements, although the adjacency matrix is based on the basic structure of the human body and remains unchanged, the angle matrix will vary greatly. For example, in a spinning movement, the angles of various parts of the body change continuously with the spin, and these angle changes are accurately recorded by the angle matrix.
[0134] Determining the adjacency matrix is based on the anatomical structure knowledge of the human body, and this structural knowledge is pre-coded into the algorithm. Through simple logical judgments, the adjacency matrix is constructed according to the connection relationships of each key point.
[0135] For the calculation of the angle matrix, it is based on geometric principles. In a two-dimensional image coordinate system, for the coordinates of any two key points, the angle between their connecting line and the horizontal direction (x-axis) can be calculated through trigonometric functions. By performing such calculations for each pair of connected key points, the angle matrix can be constructed.
[0136] Still taking the moment of takeoff of a high jumper as an example, the server calculates the structural features of the action to be measured based on the preset key point weights, the adjacency matrix, and the angle matrix.
[0137] The preset key point weights are determined according to the importance of different key points in describing the action structure. For example, in the high jump action, the weights of the ankles, knees, and waist may be relatively high because these parts play a key role in power transmission and body posture adjustment during takeoff; while the weights of the wrists and elbows may be relatively low.
[0138] The principle of this calculation process lies in quantifying the structural information of human postures. The adjacency matrix represents the connection relationships between key points, determining which parts of the angle information are relevant; the angle matrix provides the angle information of each connected part, reflecting the relative directions of various body parts; and the preset key point weights are weighted according to the importance of different key points in the action.
[0139] By calculating these three factors according to a formula, the geometric and physiological relationships of the human action structure are transformed into a quantifiable feature vector. This feature vector can mathematically represent the structural features of the action and can be used for subsequent comparison and matching with the structural features of the reference action, thereby realizing action recognition.
[0140] In an embodiment of the present invention, the structural features of the to-be-detected action are calculated based on the preset key point weights, the adjacency matrix, and the angle matrix, and can be implemented through the following examples.
[0141] According to the formula: The structural features of the to-be-detected action are calculated;
[0142] Wherein, is the structural feature of the to-be-detected action, is the adjacency matrix, and the adjacency matrix is used to represent the relationships between human key points, is the angle matrix used to represent the horizontal angles of the connections between
[0143] In an embodiment of the present invention, for example, it is assumed that the server is processing video surveillance data of a dance performance, and the purpose is to identify different dance movements of the dancer.
[0144] In the dance performance scenario, the dancer makes various complex movement postures. After the server processes the human body position area of the dancer through a human pose estimation model, the position information of multiple human key points is obtained, and then the adjacency matrix, the angle matrix, and the preset key point weights are determined.
[0145] For example, for a specific dance movement, the human key points include the head, neck, shoulders, elbows, wrists, waist, knees, and ankles, etc. The adjacency matrix determines the connection relationships between these key points according to the physiological structure of the human body. For example, the head is connected to the neck, the neck is connected to the shoulders, etc., and the corresponding matrix elements are 1, and the unconnected elements are 0.
[0146] The angle matrix accurately reflects the horizontal angles of the lines connecting each key point. In dance movements, these angles change a lot at different times. For example, when a dancer stretches his arms, the horizontal angles of the lines connecting the shoulder and elbow, and the lines connecting the elbow and wrist will change significantly; when a dancer bends his knees, the horizontal angles of the lines connecting the waist and knees, and the lines connecting the knees and ankles will also change.
[0147] The preset key point weights are set according to the characteristics of dance movements. In dance, the movements of the waist, shoulders and knees often have a greater impact on the entire dance posture, so the key point weights of these parts are relatively high.
[0148] In this way, the server can calculate the corresponding action structure features to be tested for each action posture in the dance performance, thereby providing a basis for subsequent comparison with the benchmark action structure features and action recognition. In different dance styles, such as classical dance, modern dance or folk dance, due to the differences in posture, rhythm and style of the action, the calculated action structure features will be significantly different, which helps to accurately identify and distinguish different types of dance movements.
[0149] In an embodiment of the present invention, feature matching is performed on the semantic features of the action to be tested and the structural features of the action to be tested based on a plurality of the benchmark action semantic features and a plurality of the benchmark action structural features, a target benchmark picture with the highest matching degree is determined, and the target benchmark action category corresponding to the target benchmark picture is used as the action recognition result of the picture to be tested, which can be implemented through the following examples.
[0150] Calculating the semantic cosine similarity between the semantic feature of the action to be tested and the semantic feature of each benchmark action;
[0151] Calculating the structural cosine similarity between the action structural feature to be tested and each of the reference action structural features;
[0152] Determine whether there is a pending reference picture whose semantic cosine similarity and structural cosine similarity are both greater than a preset cosine similarity threshold;
[0153] If it exists, taking the undetermined reference picture as the target reference picture, and taking the target reference action category corresponding to the target reference picture as the action recognition result of the picture to be tested;
[0154] If not, the action recognition result of the image to be tested is determined as an unknown action category.
[0155] In an embodiment of the present invention, by way of example, assume that the server is processing image data in a mall monitoring system to identify various actions of customers. The server has obtained the semantic features of the action to be measured through previous steps and retrieved multiple benchmark action semantic features from the benchmark feature library.
[0156] For example, the action to be measured is a customer reaching for a product in front of a mall shelf, and its corresponding semantic features of the action to be measured are a high-dimensional vector that contains various information related to the semantics of this action, such as the smoothness of the action, the intention of the arm extension, etc.
[0157] The benchmark action semantic features in the benchmark feature library cover various common actions in the mall, such as the semantic features corresponding to "walking", "selecting products", "queuing for payment", etc. For the semantic features of the benchmark action of "selecting products", it is also a high-dimensional vector that can contain semantic information such as the degree of concentration of the customer's attention when selecting products and the frequency of hand movements.
[0158] The server calculates the semantic cosine similarity between the semantic features of the action to be measured and the semantic features of the benchmark action of "selecting products".
[0159] During the calculation process, the server multiplies the corresponding elements of the two vectors and accumulates them to obtain the dot product of the vectors, then calculates the norms of the two vectors respectively, and finally calculates the semantic cosine similarity through a formula. For example, if the calculated semantic cosine similarity is 0.8, this indicates that the action to be measured has a high semantic similarity to the benchmark action of "selecting products".
[0160] The server will perform such calculations for each benchmark action semantic feature in the benchmark feature library, such as calculating the semantic cosine similarity with the semantic features of actions such as "walking" and "queuing for payment".
[0161] Taking the monitoring of a sports event as an example, assume that the server is analyzing the actions of athletes. The action to be measured is a basketball player shooting a basket, and the obtained structural features of the action to be measured contain the structural relationships of various parts of the athlete's body at the moment of shooting, such as the extension angle of the arm, the inclination of the body, etc.
[0162] The benchmark feature library contains the benchmark action structural features of various sports actions, such as the structural features of actions such as "dribbling", "passing the ball", and "defending". For the structural features of the benchmark action of "shooting a basket", it is also a vector that contains the structural relationships of various parts of the body in the standard posture of the shooting action.
[0163] The server calculates the structural cosine similarity between the structural features of the action to be measured and the structural features of the reference action of "shooting a basket". Similarly, according to the cosine similarity formula, the similarity degree between two vectors in the structural feature space is calculated. Suppose the calculated structural cosine similarity is 0.75, which indicates that structurally, the action to be measured has a certain similarity with the reference action of "shooting a basket".
[0164] Just like calculating the semantic cosine similarity, the server will perform such calculations for each structural feature of the reference action.
[0165] In the airport security check scenario, the server has calculated all the semantic cosine similarities and structural cosine similarities. Suppose the preset cosine similarity threshold is 0.7.
[0166] For the action of a passenger undergoing security check, the server compares it with various security check-related actions (such as "raising the hand for inspection", "turning around", etc.) in the reference feature library. If there is a certain reference action, for example, "raising the hand for inspection", whose semantic cosine similarity is 0.8 and structural cosine similarity is 0.72, and both of these similarities are greater than the preset cosine similarity threshold of 0.7, then the reference picture containing the action of "raising the hand for inspection" is determined as the pending reference picture.
[0167] If for all the reference actions, there is no such situation, that is, no reference action has both semantic cosine similarity and structural cosine similarity greater than the preset threshold, then it means that no action with a high enough matching degree with the action to be measured is found in the existing reference feature library.
[0168] On the factory production line, the server determines whether there is a pending reference picture through the previous steps. If there is, for example, when identifying the action of a worker operating a certain device, it is found that the reference picture corresponding to the reference action of "operating the device button" is the pending reference picture, then this pending reference picture is determined as the target reference picture, and the reference action category of "operating the device button" is used as the recognition result of the worker's action in the picture to be measured.
[0169] If there is no pending reference picture, such as when detecting a new special action not covered in the reference feature library (which can be a unique operating action of a worker temporarily), then the server determines the recognition result of the action in this picture to be measured as an unknown action category, which helps with the subsequent analysis of new actions and the update of the reference feature library.
[0170] In the embodiments of the present invention, the following implementation manners are also provided.
[0171] In the case where there are multiple pending reference pictures, calculate the Euclidean distances between the pending action semantic features and the pending action structure features corresponding to each pending reference picture and the semantic features and the structure features of the to-be-tested action.
[0172] According to the calculation results of the Euclidean distances, determine the pending reference picture with the highest matching degree as the target reference picture.
[0173] In an embodiment of the present invention, by way of example, it is assumed that the server is responsible for identifying the actions of patients and medical staff in a hospital ward to ensure safety and normal nursing procedures in the ward.
[0174] When analyzing the actions in the ward, the server has completed the previous steps, that is, calculating the semantic cosine similarity between the semantic features of the to-be-tested action and the semantic features of each reference action, and the structural cosine similarity between the structural features of the to-be-tested action and the structural features of each reference action, and determining multiple pending reference pictures according to a preset cosine similarity threshold.
[0175] For example, the to-be-tested action is an action of a nurse beside a hospital bed, and this action may include elements such as leaning forward, bending the arm, and some operating actions of the hand. After the previous similarity calculation, it is found that the reference pictures corresponding to three reference actions have become pending reference pictures, namely "taking the patient's temperature", "adjusting the infusion rate", and "helping the patient turn over". These three actions have a high similarity with the to-be-tested action both semantically and structurally, so the reference pictures corresponding to them are determined as pending reference pictures.
[0176] For the pending reference picture of "taking the patient's temperature", the server obtains the corresponding pending action semantic features and pending action structure features. The pending action semantic features may include semantic information such as the degree of concentration and the gentleness of the action when the nurse takes the temperature; the pending action structure features include the positional relationship of each part of the nurse's body in this action, such as the relative position of the arm and the patient's body, and the bending angle of the body.
[0177] For the pending reference picture of "adjusting the infusion rate", repeat the above process. Its pending action semantic features include semantic information such as the nurse checking the infusion device and operating the adjustment device; the pending action structure features include structural information such as the relative position of the nurse's hand and the infusion device and the body posture.
[0178] Calculate the Euclidean distances of the semantic features and the structural features respectively according to the Euclidean distance formula, and then add them up to obtain the total Euclidean distance.
[0179] For the pending reference picture of "helping the patient turn over", the pending action semantic features include semantic information such as the collaborative information between the nurse and the patient, the strength of the action, etc.; the pending action structural features include the positional relationships of various parts of the nurse's body during the process of assisting in turning over, etc.
[0180] Similarly, calculate the Euclidean distances of the semantic features and the structural features and add them up to obtain the total Euclidean distance.
[0181] After comparison by the server, according to the magnitudes of the corresponding Euclidean distances, it is determined that the overall difference between the pending action corresponding to the pending reference picture of "adjusting the infusion rate" and the action to be measured is the smallest, that is, the matching degree is the highest.
[0182] Then, the server determines the pending reference picture of "adjusting the infusion rate" as the target reference picture, and takes the reference action category of "adjusting the infusion rate" as the final recognition result of the nurse's action in the picture to be measured. In this way, through the calculation of the Euclidean distance, in the case of multiple pending reference pictures, the reference action category that best matches the action to be measured can be more accurately determined, improving the accuracy of action recognition.
[0183] The Euclidean distance is a method for measuring the distance between two points (here are feature vectors) in a multi-dimensional space. In the scenario of action recognition, both the semantic feature vector and the structural feature vector represent the feature representations of the action in different dimensions. By calculating the Euclidean distance, the degree of difference in both the semantic and structural aspects can be comprehensively considered.
[0184] When there are multiple pending reference pictures, the Euclidean distances between the pending action features corresponding to each pending reference picture and the action features to be measured reflect their overall differences. The smaller the distance, the closer they are to the action to be measured in terms of semantics and structure. By comparing these Euclidean distances and selecting the pending reference picture corresponding to the smallest distance as the target reference picture, the most matching action category can be selected from multiple similar candidate actions, thereby improving the accuracy and reliability of the entire action recognition system.
[0185] In the embodiment of the present invention, the obtaining of the picture to be measured can be implemented through the following examples.
[0186] Obtain the picture to be measured and the business scenario corresponding to the picture to be measured;
[0187] The obtaining of the reference action semantic features and the reference action structural features corresponding to multiple reference pictures from the pre-set reference feature library includes;
[0188] Obtain the reference action semantic features and the reference action structural features corresponding to multiple reference pictures from the pre-set reference feature library according to the business scenario.
[0189] In an embodiment of the present invention, by way of example, in an intelligent transportation system, the server is responsible for recognizing the actions of pedestrians, vehicle drivers, etc. on the road.
[0190] The server obtains the pictures to be tested from the traffic monitoring camera network. These cameras are distributed at various intersections, main roads, parking lots, etc. in the city. At the same time, the server obtains the business scenario information corresponding to these pictures to be tested.
[0191] For example, for the picture to be tested taken by a monitoring camera at an intersection, the corresponding business scenario is "Intersection Traffic Supervision". The pictures to be tested in this scenario include images of pedestrians crossing the road, vehicles turning, traffic police directing traffic, etc.
[0192] For another example, for the picture to be tested taken by a camera at the entrance of a parking lot, the corresponding business scenario is "Parking Lot Vehicle and Personnel In-and-Out Management", and the picture may include actions such as a driver parking and taking a card, and a pedestrian guiding a vehicle.
[0193] The server obtains these pictures to be tested and the corresponding business scenario information in order to perform more accurate action recognition subsequently. Because in different business scenarios, the types of actions to be concerned about and their characteristics vary greatly. In the intersection scenario, the key actions to be concerned about are those related to compliance with traffic rules, such as whether pedestrians cross the road according to the traffic lights and whether vehicles turn correctly; while in the parking lot scenario, more attention is paid to actions related to parking operations and personnel guidance.
[0194] The server receives the pictures to be tested from the internal monitoring cameras of a bank, and the corresponding business scenario is "Bank Business Hall Security and Service Monitoring".
[0195] A large number of reference pictures in different business scenarios and their corresponding reference action semantic features and reference action structure features are stored in the reference feature library. For the business scenario of the bank business hall, the server obtains the features of the reference pictures related thereto from the reference feature library.
[0196] For example, the reference pictures may include pictures of tellers handling business for customers (such as actions of counting money, stamping, etc.), customers operating at self-service terminals (such as actions of withdrawing money, querying, etc.), security guards patrolling (such as actions of walking, observing the surrounding environment, etc.).
[0197] For the reference pictures of tellers handling business for customers, the reference action semantic features contain semantic information related to the business operation process. For example, the semantic features of the money-counting action may include the speed of money-counting, concentration, etc., and the semantic features of the stamping action may include the accuracy of the stamping position, the standardization of the action, etc.
[0198] Its reference action structure features include the positional relationships of various parts of the teller's body during the operation. For example, the movement trajectory of the arm during banknote counting, the relative positions of the hand, the banknote counting machine, and the banknotes, the swinging amplitude of the arm during stamping, the degree of forward leaning of the body, etc.
[0199] For the reference pictures of customers' operations on self-service terminals, the reference action semantic features include semantic information such as the customers' operation intentions and proficiency levels. For example, the degree of caution when entering the password during a withdrawal operation, the decisiveness when selecting the withdrawal amount, etc.
[0200] The reference action structure features include the relative positions of the customers' bodies and the self-service terminals, the action postures of the hands on the operation screens and keyboards, etc.
[0201] For the reference pictures of security guards' patrols, the reference action semantic features include semantic information such as the security guards' alertness and the regularity of the patrol routes.
[0202] The reference action structure features include the stride lengths of the security guards' walks, the angles of body rotation to observe the surrounding environment, etc.
[0203] By obtaining the relevant reference action semantic features and reference action structure features from the reference feature library according to the business scenarios, the server can identify the actions in the pictures to be tested more pertinently. For example, in the scenario of a bank business hall, when the server identifies the actions of the tellers in the pictures to be tested, the reference action features related to the tellers' handling of business can provide accurate references, improving the accuracy and efficiency of action recognition. Similarly, the action recognition of customers and security guards can also be carried out more effectively within the framework of this specific business scenario.
[0204] To be able to more clearly describe the solution provided by the embodiments of the present invention, a relatively complete implementation manner is provided below.
[0205] The present invention proposes a general action recognition method based on semantic feature and structure feature matching. This method realizes a general feature extraction method, that is, first uses the YOLOv5 model for human detection, locates the human body area and crops it, and then uses the action autoencoder model AAEM to extract the semantic feature S1 of the human body behavior action; then uses the AlphaPose model for human key point detection to calculate the structure feature S2 of the human key points; based on this feature extraction method, the present invention needs to extract key frames from the action video as reference pictures, and extract the semantic features and structure features of the reference pictures, and save them in the feature library. In practical applications, only need to use the same method to extract the semantic features and structure features of the picture to be tested, and then perform feature matching in the feature library of the reference pictures to obtain the matching result. The implementation method technology of the present invention relates to the fields of deep learning, image processing, and computer vision technology, and can be applied to fields such as video surveillance, human-computer interaction, and video behavior analysis.
[0206] To achieve the above object, the present invention adopts the following technical solutions, mainly including 4 stages. Please refer to Figure 2 。
[0207] The first stage is human detection, using the pre-trained YOLOv5 model for human detection to detect and locate the regional position of the human body in the image;
[0208] The second stage is to build and train an action autoencoder model AAEM to extract the semantic feature S1 of the human body behavior action;
[0209] The third stage is human key point detection, using the pre-trained AlphaPose model for human key point detection to calculate the structure feature S2 of the human key points;
[0210] The fourth stage is to build a reference picture dataset of action key frames; the reference picture is used as the reference image and template for reference action recognition, and it is necessary to extract the semantic feature B S1 and the structure feature B S2 , and save them in the feature library of the system.
[0211] The fifth stage is to achieve feature matching, extract the semantic feature T S1 and the structure feature T S2 of the test picture, and then perform feature matching in the feature library of the reference pictures, and select the most similar reference picture as the matching result.
[0212] The first stage is human detection: As an important application in the field of computer vision, human detection aims to accurately identify and locate the human body position area in images or videos. This technology is not only crucial for fields such as security monitoring, intelligent security, and human-computer interaction, but also the foundation for modern intelligent systems to achieve efficient action recognition and behavior analysis. To achieve this goal, this solution directly adopts the pre-trained model YOLOv5, which has been trained for object detection tasks on the COCO (Common Objects in Context) dataset, covering various complex scenarios and multiple object categories, including the human body. By using the YOLOv5 model, we can quickly and accurately identify the human body in the image and precisely locate its position area. This provides a solid foundation for subsequent action recognition, behavior analysis, and other work.
[0213] The second stage is to build and train an action autoencoder model AAEM for extracting the semantic feature S1 of human behavior actions. The action autoencoder model consists of an encoder and a decoder. The encoder is used to map the input action picture data to a low-dimensional semantic feature space, and the decoder is used to reconstruct the representation in the low-dimensional semantic feature space back to the original action picture data. Please refer to Figure 3 and Figure 4 .
[0214] Build an Action AutoEncoder Model, abbreviated as AAEM. The encoder consists of 5 downsampling modules (Blocks), which map the input RGB image of 224×224 to a low-dimensional semantic feature space and encode it into a feature vector of 512. The decoder consists of 5 upsampling modules (upBlocks), which reconstruct the 512 feature vector back to the original image. The Block module adopts the Conv (convolution)-BN-PRelu-Conv-BN-PRelu structure, while the upBlock module adopts the DeConv (transposed convolution)-BN-PRelu-DeConv-BN-PRelu structure. These two structures can not only effectively extract the key features in the image and realize the functions of upsampling and downsampling, but also introduce the non-linear activation function PReLU. This activation function has a learnable parameter w, which can adaptively learn and correct the parameters of the linear unit, improve the fitting ability of the model, and reduce the risk of overfitting, enabling the model to learn more complex mapping relationships. In addition, both the Block and upBlock modules add a channel attention module, which assigns different importance weights to different channels, enabling the model to focus on the features that are more critical to the current task, thereby enhancing the overall feature representation ability of the model. The channel attention module is mainly composed of the Global Pooling-FC-sigmoid structure, and the specific implementation is as follows:
[0215] Global Pooling is responsible for performing global pooling operations on the feature maps of each channel, thereby obtaining a one-dimensional vector. This vector contains the feature information of all positions in the channel and is crucial for the subsequent learning process.
[0216] The fully connected layer (FC) performs a fully connected operation on the vector output by the Global Pooling layer and learns the weights of each channel through non-linear transformation. These weights reflect the importance of different channels to the current task and are the core of the attention mechanism.
[0217] The normalization layer normalizes the weights output by the FC layer (such as the sigmoid function), making the sum of the weights of all channels equal to 1. In this way, the weight of each channel becomes a probability value between 0 and 1, that is, the obtained is the attention coefficient of each channel, which is convenient for the subsequent network layer to perform weighted operations.
[0218] The attention application multiplies the obtained channel attention coefficient vector with the original feature map element-wise, thereby weighting the features of each channel. The weighted feature map retains the original spatial structure, but the influence of each of its channels has been adaptively adjusted.
[0219] Training the AAEM model: The AAEM model belongs to the autoencoder model and is an unsupervised learning-based method for data dimension compression and feature representation. Without the need to label image categories, the data annotation cost of the AAEM model is almost zero, and the model can be directly trained using the images of existing open-source data. In this solution, the open-source dataset ImageNet is directly used. This dataset contains more than 14 million images, covering more than 20,000 categories. In this solution, ILSVRC2012_img_train of ImageNet is used as the training set, and ILSVRC2012_img_val is used as the validation set; the above-mentioned training set is used for AAEM model training and optimization; the above-mentioned test set is used for AAEM model testing and evaluating the performance of model metrics. When training the model, the input data is uniformly scaled to 224×224, the batch size batch_size = 256, the optimization algorithm uses the Adam optimizer, the initial learning rate is set to lr = 0.001, and the mean squared error loss function (MSE Loss) is adopted. The entire dataset is iterated epoch = 200 times; during the training process, after each complete data iteration, the MSE of the model on the test set is calculated once, and the model weights with the lowest current mean squared error MSE are saved; in subsequent actual use, only the model weight file and the corresponding inference program code need to be retained.
[0220] Deploying the AAEM model: After training, for the obtained action autoencoder model, in actual use, only the encoder part of this solution is required to extract image features, while the decoder does not participate in the inference calculation;
[0221] The third stage is to perform human keypoint detection. The pre-trained AlphaPose model is used for human keypoint detection to calculate the structural features S2 of human keypoints; the specific implementation steps are as follows:
[0222] For human keypoint detection, this solution directly adopts the pre-trained model AlphaPose. This model has been fully trained on the human keypoint detection task of the COCO (Common Objects in Context) dataset and supports the detection of 17 keypoints of the whole body. By using the AlphaPose model, we can predict 17 keypoints of the human body, including the position information of keypoints such as the head, hands, feet, and body. Please refer toFigure 5 。
[0223] Calculate the structural features S2 of human body key points:
[0224] ;
[0225] For the AlphaPose model trained on the COCO dataset, n = 17, that is, its output has 17 human body key points;
[0226] A is the adjacency matrix, which is used to represent the point-to-point connection relationship between human body key points. 0 means no connection, and 1 means connection; According to the schematic diagram of human body key points in the COCO dataset, when belongs to the set {(1,2),(0,1),(0,2),(1,3),(2,4),(3,5),(4,6),(15,13),(13,11),(16,14),(14,12),(11,12),(5,11),(6,12),(5,6),(5,7),(6,8),(7,9),(8,10)}, , when does not belong to this set, then 0.
[0227] D is the angle matrix, which is used to measure the horizontal angle of the connection line between any two key points;
[0228] ;
[0229] where I represents the pixel coordinates of (x, y) in the image.
[0230] The W weight matrix is used to configure the weight size between any two key points, and the default value is 1.
[0231] The fourth stage is to construct the benchmark picture dataset of action key frames; The benchmark pictures, as the reference images and templates for benchmark action recognition, need to extract the semantic feature B S1 and the structural feature B S2 , and save them in the feature library of the system.
[0232] The fifth stage is to achieve feature matching. Extract the semantic feature T S1 and the structural feature T S2 of the test picture, and then perform feature matching in the feature library of the benchmark pictures to select the most similar benchmark picture as the matching result. The specific method is as follows:
[0233] The test picture first performs YOLOv5 human detection, crops the human body area, and then uses the AAEM encoder to extract the semantic feature T S1, then use the AlphaPose model to obtain 17 key points of the human body and calculate the structural feature T of the human body key points S2 .
[0234] Calculate the cosine similarity of semantic features between the test picture and the reference picture in the feature library and the cosine similarity of structural features .
[0235] ;
[0236] When the cosine similarity of semantic features and the cosine similarity of structural features are respectively greater than the confidence thresholds T1 and T2, the action category corresponding to the reference picture is the test picture matching result; otherwise, it belongs to an unknown category.
[0237]
[0238] Please refer to Figure 6 , Figure 6 A general action recognition device 110 based on semantic feature and structural feature matching provided by an embodiment of the present invention, including:
[0239] An acquisition module 1101, configured to acquire a to-be-tested picture, input the to-be-tested picture into a pre-trained human body recognition model to obtain a human body position area; call a pre-trained action autoencoder model to perform feature extraction on the human body position area to obtain a to-be-tested action semantic feature; call a pre-trained human body pose estimation model to perform feature extraction on the human body position area to obtain a to-be-tested action structural feature; obtain the reference action semantic features and reference action structural features corresponding to a plurality of reference pictures from a pre-set reference feature library, and each of the reference pictures is marked with a corresponding reference action category;
[0240] An identification module 1102, configured to perform feature matching on the to-be-tested action semantic feature and the to-be-tested action structural feature according to the plurality of reference action semantic features and the plurality of reference action structural features, determine a target reference picture with the highest matching degree, and use the target reference action category corresponding to the target reference picture as the action recognition result of the to-be-tested picture.
[0241] It should be noted that the implementation principle of the foregoing general action recognition device 110 based on semantic feature and structural feature matching can refer to the implementation principle of the foregoing general action recognition method based on semantic feature and structural feature matching, which will not be elaborated here. It should be understood that the division of each module of the above device is only a division of logical functions. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by processing elements; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by processing elements, and some modules can be implemented in the form of hardware. For example, the general action recognition device 110 based on semantic feature and structural feature matching can be a separately established processing element, or can be integrated in a certain chip of the above device. In addition, it can also be stored in the memory of the above device in the form of program code, and the function of the above general action recognition device 110 based on semantic feature and structural feature matching can be called and executed by a certain processing element of the above device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. The processing element described here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.
[0242] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more application specific integrated circuits (ASICs), or, one or more digital signal processors (DSPs), or, one or more field programmable gate arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element scheduling program code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processors that can call program code. Again, these modules can be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0243] An embodiment of the present invention provides a computer device 100. The computer device 100 includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the foregoing general action recognition device 110 based on semantic feature and structural feature matching. As Figure 7 shown, Figure 7The computer device 100 provided in the embodiment of the present invention is a structural block diagram. The computer device 100 includes a general action recognition device 110 based on semantic feature and structural feature matching, a memory 111, a processor 112 and a communication unit 113.
[0244] In order to realize data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, the electrical connection between these elements can be realized through one or more communication buses or signal lines. The general action recognition device 110 based on semantic feature and structural feature matching includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the general action recognition device 110 based on semantic feature and structural feature matching stored in the memory 111, such as the software function modules and computer programs included in the general action recognition device 110 based on semantic feature and structural feature matching.
[0245] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the aforementioned general action recognition device 110 based on matching of semantic features and structural features.
[0246] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.
Claims
1. A general action recognition method based on semantic feature and structure feature matching, characterized in that Including: Obtain a to-be-tested picture, and input the to-be-tested picture into a pre-trained human body recognition model to obtain a human body position area; Call a pre-trained action auto-encoding model to extract features from the human body position area to obtain to-be-tested action semantic features; Call a pre-trained human body pose estimation model to extract features from the human body position area to obtain to-be-tested action structure features; Obtain the benchmark action semantic features and benchmark action structure features corresponding to a plurality of benchmark pictures from a pre-set benchmark feature library, and each of the benchmark pictures is marked with a corresponding benchmark action category; According to the plurality of benchmark action semantic features and the plurality of benchmark action structure features, perform feature matching on the to-be-tested action semantic features and the to-be-tested action structure features, determine a target benchmark picture with the highest matching degree, and use the target benchmark action category corresponding to the target benchmark picture as the action recognition result of the to-be-tested picture; The step of calling a pre-trained human body pose estimation model to extract features from the human body position area to obtain to-be-tested action structure features includes: Call a pre-trained human body pose estimation model to perform key point detection on the human body position area to obtain the position information of multiple human body key points; According to the position information of the multiple human body key points, determine a corresponding adjacency matrix and angle matrix; According to a preset key point weight, the adjacency matrix, and the angle matrix, calculate to obtain the to-be-tested action structure features; The step of calculating the to-be-tested action structure features according to a preset key point weight, the adjacency matrix, and the angle matrix includes: According to the formula: the structural features of the action to be measured are calculated; Among them, is the structural feature of the action to be measured, is the adjacency matrix, and the adjacency matrix is used to represent the relationship between individual human body key points, is the angle matrix for representing the horizontal angle of the line connecting two individual human body key points, is the preset key point weight.
2. The method according to claim 1, characterized in that, The action auto-encoding model is obtained by the following method, including: Construct an initial action auto-encoding model, the initial action auto-encoding model includes five cascaded downsampling components and five upsampling components, the downsampling component is a Conv-BN-PRelu-Conv-BN-PRelu structure, the upsampling component is a DeConv-BN-PRelu-DeConv-BN-PRelu structure, and each of the downsampling components and each of the upsampling components are configured with a channel attention component, and the channel attention component includes a global pooling layer, a fully connected layer, and a normalization layer; Train the initial action auto-encoding model based on an unsupervised mechanism until a preset mean square error loss function training termination condition is reached to obtain the action auto-encoding model.
3. The method according to claim 1, wherein The step of performing feature matching on the to-be-tested action semantic features and the to-be-tested action structure features according to the plurality of benchmark action semantic features and the plurality of benchmark action structure features, determining a target benchmark picture with the highest matching degree, and using the target benchmark action category corresponding to the target benchmark picture as the action recognition result of the to-be-tested picture includes: Calculate the semantic cosine similarity between the to-be-tested action semantic features and each of the benchmark action semantic features; Calculate the structural cosine similarity between the to-be-tested action structure features and each of the benchmark action structure features; Judge whether there is a pending benchmark picture whose semantic cosine similarity and structural cosine similarity are both greater than a preset cosine similarity threshold; If it exists, use the to-be-determined reference picture as the target reference picture, and use the target reference action category corresponding to the target reference picture as the action recognition result of the to-be-tested picture; If it does not exist, determine the action recognition result of the to-be-tested picture as an unknown action category.
4. The method according to claim 3, wherein The method further includes: In the case where there are multiple to-be-determined reference pictures, calculate the Euclidean distances between the to-be-determined action semantic features and the to-be-determined action structure features corresponding to each to-be-determined reference picture and the to-be-tested action semantic features and the to-be-tested action structure features; According to the calculation results of the Euclidean distances, determine the to-be-determined reference picture with the highest matching degree as the target reference picture.
5. The method according to claim 1, wherein The obtaining the to-be-tested picture includes: Obtain the to-be-tested picture and the service scenario corresponding to the to-be-tested picture; The obtaining the reference action semantic features and the reference action structure features corresponding to multiple reference pictures from a pre-set reference feature library includes; According to the service scenario, obtain the reference action semantic features and the reference action structure features corresponding to multiple reference pictures from a pre-set reference feature library.
6. A general action recognition device based on semantic feature and structure feature matching, characterized in that, It includes: An obtaining module, configured to obtain a to-be-tested picture, input the to-be-tested picture into a pre-trained human body recognition model to obtain a human body position area; call a pre-trained action auto-encoder model to extract features from the human body position area to obtain to-be-tested action semantic features; Call a pre-trained human body pose estimation model to extract features from the human body position area to obtain to-be-tested action structure features; Obtain the reference action semantic features and the reference action structure features corresponding to multiple reference pictures from a pre-set reference feature library, and each of the reference pictures is marked with a corresponding reference action category; An identification module, configured to perform feature matching on the to-be-tested action semantic features and the to-be-tested action structure features according to the multiple reference action semantic features and the multiple reference action structure features, determine the target reference picture with the highest matching degree, and use the target reference action category corresponding to the target reference picture as the action recognition result of the to-be-tested picture; The obtaining module is specifically configured to: Call a pre-trained human body pose estimation model to perform key point detection on the human body position area to obtain multiple human body key point position information; according to the multiple human body key point position information, determine the corresponding adjacency matrix and angle matrix; according to a preset key point weight, the adjacency matrix, and the angle matrix, calculate to obtain the to-be-tested action structure features; The obtaining module is further specifically configured to: According to the formula: The structural features of the action to be measured are calculated; Among them, is the structural feature of the action to be measured, is the adjacency matrix, and the adjacency matrix is used to represent the relationship between individual human body key points, is the angle matrix used to represent the horizontal angle of the connection line between two individual human body key points, is the preset key point weight.
7. A computer device, characterized in that, The computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1-5.
8. A readable storage medium, characterized in that, The readable storage medium includes a computer program. When the computer program runs, it controls the computer device where the readable storage medium is located to execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
User running posture analysis method and device, treadmill and storage medium
CN115188062A
Pedestrian crossing intention prediction method based on pedestrian crossing intention prediction network
CN117557979A