Person Activity Analysis Method and System Based on Image Recognition

Through the convolutional network and regional screening network, the target entities in the video stream are extracted and segmented, combined with the cascading pyramid network and multi-scale processing technology, human body posture estimation and motion feature analysis are carried out, which solves the problem of misjudgment of action recognition in the existing technology and realizes accurate identification and classification of human body activities.

CN118212688BActive Publication Date: 2025-06-03HASO SOFT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410297224.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-15
Publication Date
2025-06-03
Estimated Expiration
2044-03-15

AI Technical Summary

Technical Problem

In the prior art, if action recognition is completed only through feature alignment, there may be a risk of behavioral misjudgment.

Method used

The features of image frames are extracted through the convolution network, combined with the regional screening network to determine the candidate areas of the target entity, used the full connection layer and the full convolution layer to determine the contour of the target entity and segmented it. The foot position is further determined through the cascading pyramid network, calculated the vertical projection, generated the target image, and performed multi-scale processing to estimate the human body posture, mapped to three-dimensional space for time series analysis, obtain the action features, and constructed a linear kernel function for action classification based on the action analysis model.

Benefits of technology

It improves the accurate identification of human activities, reduces the risk of misjudgment of behavior, and achieves a more detailed understanding and classification of action characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118212688B_ABST
    Figure CN118212688B_ABST
Patent Text Reader

Abstract

The present invention provides a method and system for analyzing human activities based on image recognition, which relates to the technical field of image recognition and includes: obtaining a video stream, extracting image frames, extracting features in the image frames and generating a convolutional feature map, determining candidate regions of target entities through a region screening network, combining the context information corresponding to the candidate regions to obtain candidate positions, determining the contours of the target entities and performing segmentation to obtain segmentation contours; based on the segmentation contours, determining the foot positions of the target entities in the bounding boxes, calculating the vertical projections of the target entities, generating target images, performing human pose estimation on the target images to obtain two-dimensional human poses, mapping the two-dimensional human poses to the three-dimensional space to obtain three-dimensional human poses, and performing time series analysis on the three-dimensional human poses to obtain action features; based on the action features, constructing a linear kernel function, and based on the linear kernel function, combining motion action classification to determine the human activities of the target entities.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular, to a method and system for analyzing human activities based on image recognition. Background Art

[0002] In the prior art, CN111523361A discloses a human behavior recognition method. This method first extracts information of two modalities, namely a first image representing static information and a second image representing dynamic information, from video data. Then, an implicit alignment is performed on the first image and the second image information using a convolutional neural network with an attention mechanism. The implicitly aligned features are further mapped into a common subspace for explicit alignment. Then, a sparse shrinkage deep autoencoder is used to perform deep fusion on the aligned different-modal features. Finally, a deep belief network is trained with the highly robust and strongly discriminative features obtained after fusion to achieve the function of high-precision human behavior recognition.

[0003] In summary, although the prior art can perform image recognition and analysis through a convolutional neural network, only action recognition is completed through feature alignment, which may pose a risk of misjudgment of behaviors. Therefore, a solution is needed to solve the problems existing in the prior art. Summary of the Invention

[0004] The embodiments of the present invention provide a method and system for analyzing human activities based on image recognition, which can at least solve some problems existing in the prior art.

[0005] In the first aspect of the embodiments of the present invention, a method for analyzing human activities based on image recognition is provided, including:

[0006] Obtain a video stream, extract image frames from the video stream, extract features in the image frames through a convolutional network and generate a convolutional feature map, determine candidate regions of a target entity through a region screening network for the convolutional feature map, combine the context information corresponding to the candidate regions to obtain candidate positions, determine the contour of the target entity through a fully connected layer and a fully convolutional layer and perform segmentation to obtain the segmentation contour of the target entity;

[0007] Based on the segmentation contour, determine the foot position of the target entity in a bounding box through a cascaded pyramid network, calculate the vertical projection of the target entity based on the principle of perspective projection to generate a target image, perform human pose estimation on the target image through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human pose, map the two-dimensional human pose to a three-dimensional space to obtain a three-dimensional human pose, and perform time series analysis on the three-dimensional human pose to obtain action features;

[0008] Based on the action features, a linear kernel function is constructed in combination with a preset action analysis model. Based on the linear kernel function and combined with motion action classification, the human activities of the target entity are determined.

[0009] In an alternative embodiment,

[0010] The video stream is acquired, the image frames in the video stream are extracted, the features in the image frames are extracted through a convolutional network to generate a convolutional feature map, and the candidate regions of the target entity are determined through a region screening network. Combining the context information corresponding to the candidate regions, the candidate positions obtained include:

[0011] An input video stream is acquired through a pre-set video device, images are extracted frame by frame in the video stream through a video processing library to obtain a plurality of image frames, the image frames are input into a pre-set convolutional neural network, and the image frames are propagated forward in the convolutional neural network. Each convolutional layer extracts features from the image frames to generate corresponding convolutional feature maps;

[0012] Based on the convolutional feature map, the region screening network proposes the bounding boxes and corresponding scores of candidate regions at each position through an anchor box mechanism, arranges the candidate regions in descending order according to the score values of each candidate region, selects the candidate regions in the top 10%, determines the coordinate information for each selected candidate region, and extracts the corresponding image regions in the image frames;

[0013] For each selected candidate region, based on a preset context information encoding module, the features of the surrounding regions of each candidate region are determined by expanding the receptive field, the context information of each candidate region is determined, and the candidate positions are determined in combination with the position information of the candidate regions.

[0014] In an alternative embodiment,

[0015] The contour of the target entity is determined through a fully connected layer and a fully convolutional layer and segmented to obtain the segmentation contour of the target entity, including:

[0016] Based on the candidate positions and the convolutional feature map, the fully connected layer and the fully convolutional layer are initialized, and the convolutional feature map is added to the input interface of the fully connected layer;

[0017] The fully connected layer maps the convolutional feature map to a high-dimensional space through feature mapping and obtains the high-level semantic information corresponding to the convolutional feature map;

[0018] The obtained high-level semantic information is added to the fully convolutional layer, the high-level semantic information is restored to the original resolution through an upsampling operation, and the initial segmentation contour of the target entity is generated through an activation function.

[0019] Fill and smooth the initial segmentation contour, compare the processed initial segmentation contour with the original image and evaluate the confidence. If the confidence is greater than the preset confidence threshold, add the initial segmentation contour as the segmentation contour of the target entity to the segmentation contour set and save it.

[0020] In an alternative embodiment,

[0021] Based on the segmentation contour, determine the foot position of the target entity in the bounding box through a cascaded pyramid network, calculate the vertical projection of the target entity based on the perspective projection principle, and generate a target image, including:

[0022] Extract the image features of the image frame based on the convolutional feature map, construct multiple convolutional layers and upsampling layers, generate multiple first feature maps with different resolutions by stacking convolutional kernels, and form a pyramid network. In the cascaded pyramid network, set multiple cascaded detection modules;

[0023] Based on the cascaded detection modules, detect the feet of the first feature map in the order of increasing resolution. First, obtain the approximate position of the feet of the target entity through the first feature map with a lower resolution, and refine the position of the feet by detecting the first feature map with a higher resolution;

[0024] For each detected foot region, adjust the bounding box of the detected foot region through the bounding box regression technique until it completely coincides with the foot image of the target entity. Based on the foot position of the target entity, calculate the position of the vertical projection point of the lowest point of the feet in the image on the ground to generate a target image.

[0025] In an alternative embodiment,

[0026] For the target image, perform human pose estimation through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human pose, map the two-dimensional human pose to the three-dimensional space to obtain a three-dimensional human pose, and perform time series analysis on the three-dimensional human pose to obtain action features, including:

[0027] For the target image, based on a pre-set multi-scale processing mechanism, extract the feature representations of the target image at multiple scales, fuse the features extracted at different scales to obtain a fused feature map;

[0028] Based on the fused feature map, learn features at different scales and determine the images corresponding to target entities of different sizes. Based on the images corresponding to the target entities, perform human pose estimation to determine the two-dimensional human pose in the images corresponding to the target entities,

[0029] Determine the internal and external camera parameters corresponding to the two-dimensional human pose, normalize the corresponding points of the contour in the two-dimensional human pose to obtain normalized coordinates, apply the camera projection model to the normalized coordinates, calculate the three-dimensional virtual coordinates corresponding to each normalized coordinate, and combine the scale factor to calculate the three-dimensional pose coordinates corresponding to the target entity to obtain the three-dimensional human pose;

[0030] Track the three-dimensional pose coordinates in consecutive frames, determine the temporal characteristics of the actions of the target entity through a sliding window, analyze the speed information and angle information corresponding to the actions based on the temporal characteristics, and determine the action characteristics.

[0031] In an alternative embodiment,

[0032] Based on the action characteristics, combine a preset action analysis model to construct a linear kernel function. Based on the linear kernel function, combine the motion action classification to determine the human activities of the target entity, including:

[0033] Generate an empty set based on the pre-acquired action characteristics and temporal characteristics, add known human actions to the empty set, and name the empty set activity information;

[0034] Based on the activity information, combine a preset action analysis model to construct a linear kernel function corresponding to the action characteristics;

[0035] Based on the linear kernel function, map the action characteristics and the temporal characteristics to a high-dimensional space to generate a high-dimensional action feature set. For each element in the high-dimensional action feature combination, combine the action analysis model to calculate the Euclidean distance between each element, and combine the position of the motion action classification in the current space to determine the action classification corresponding to each action feature, and obtain the human activities corresponding to the target entity based on the action classification.

[0036] In an alternative embodiment,

[0037] The method further includes training the action analysis model:

[0038] Collect real action data, randomly generate feature vectors according to the real action data, divide the randomly generated feature vectors into a training set and a test set, initialize the action analysis model and introduce a regularization parameter, input the feature vectors in the training set into the action analysis model to generate a first decision boundary;

[0039] Based on the first decision boundary, determine the true positive rate and false positive rate corresponding to the action analysis model, obtain the classification probability corresponding to the action analysis model, based on the classification probability, combine the true category and the model prediction to determine the confusion matrix, based on the classification probability and the confusion matrix, dynamically update the action analysis model, add the test set to the updated action analysis model to obtain the second decision boundary, and based on the second decision boundary, repeat the update of the feature vector in combination with the regularization parameter until the preset number of stop times is reached.

[0040] In a second aspect of the embodiments of the present invention, there is provided a personnel activity analysis system based on image recognition, including:

[0041] A first unit, configured to obtain a video stream, extract image frames in the video stream, extract features in the image frames through a convolutional network and generate a convolutional feature map, determine candidate regions of a target entity through a region screening network for the convolutional feature map, combine the context information corresponding to the candidate regions to obtain candidate positions, and determine the contour of the target entity through a fully connected layer and a fully convolutional layer and perform segmentation to obtain the segmentation contour of the target entity;

[0042] A second unit, configured to, based on the segmentation contour, determine the foot position of the target entity in a bounding box through a cascaded pyramid network, calculate the vertical projection of the target entity based on the perspective projection principle to generate a target image, perform human pose estimation on the target image through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human pose, map the two-dimensional human pose to a three-dimensional space to obtain a three-dimensional human pose, and perform time series analysis on the three-dimensional human pose to obtain action features;

[0043] A third unit, configured to, based on the action features, combine a preset action analysis model to construct a linear kernel function, and based on the linear kernel function, combine motion action classification to determine the human activity of the target entity.

[0044] In a third aspect of the embodiments of the present invention,

[0045] There is provided an electronic device, including:

[0046] A processor;

[0047] A memory for storing instructions executable by the processor;

[0048] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0049] In a fourth aspect of the embodiments of the present invention,

[0050] Provided is a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, the foregoing method is implemented.

[0051] In the present invention, through a convolutional network and a region screening network, the target entity in the video stream can be effectively located and segmented, and accurate target position information can be obtained. Through the time series analysis of the three-dimensional human posture, the action features of the target entity are extracted, which helps to understand the action more carefully. Combining with a preset action analysis model, a linear kernel function adapted to the action features is constructed, which improves the ability of abstracting and expressing the action features. Using the linear kernel function and motion action classification, the human activities of the target entity can be accurately classified, and the recognition and understanding of different actions can be realized. In summary, through multi-level information extraction and comprehensive analysis, the present invention realizes a comprehensive grasp of the behavior of the target entity, and provides strong technical support for more in-depth human action analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a schematic flowchart of the method for analyzing human activities based on image recognition according to an embodiment of the present invention;

[0053] Figure 2 is a schematic structural diagram of the system for analyzing human activities based on image recognition according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] The technical solutions of the present invention will be described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0056] Figure 1 is a schematic flowchart of the method for analyzing human activities based on image recognition according to an embodiment of the present invention, as Figure 1 shown, the method includes:

[0057] S1. Obtain a video stream, extract image frames from the video stream, extract features from the image frames through a convolutional network and generate a convolutional feature map, determine candidate regions of the target entity through a region screening network, combine the context information corresponding to the candidate regions to obtain candidate positions, and determine the contour of the target entity and perform segmentation through a fully connected layer and a fully convolutional layer to obtain the segmentation contour of the target entity;

[0058] The video stream refers to a continuous sequence of video images, which can come from a camera, a video file, etc. The image frame refers to a single still image in the video stream, representing an instantaneous state in the video. The convolutional feature map is used to capture different levels of abstract features in the image. The region screening network is a network used in object detection to generate candidate regions that may contain the object. The candidate region refers to a potential object region proposed by the region screening network. The target entity refers to the actual object to be detected and recognized in the video stream.

[0059] In an alternative embodiment,

[0060] The obtaining of the video stream, extracting the image frames from the video stream, extracting features from the image frames through a convolutional network and generating a convolutional feature map, determining candidate regions of the target entity through a region screening network, and combining the context information corresponding to the candidate regions to obtain candidate positions includes:

[0061] Obtain an input video stream through a pre-set video device, extract images frame by frame from the video stream through a video processing library to obtain a plurality of image frames, input the image frames into a pre-set convolutional neural network, perform forward propagation of the image frames in the convolutional neural network, and each convolutional layer extracts features from the image frames to generate corresponding convolutional feature maps;

[0062] Based on the convolutional feature map, the region screening network proposes the bounding boxes and corresponding scores of candidate regions for each position through an anchor box mechanism, sorts the candidate regions in descending order according to the score values of each candidate region, selects the candidate regions in the top 10%, determines the coordinate information for each selected candidate region, and extracts the corresponding image regions in the image frames;

[0063] For each selected candidate region, based on a pre-set context information encoding module, determine the features of the surrounding regions of each candidate region by expanding the receptive field, determine the context information of each candidate region, and combine the position information of the candidate region to determine the candidate position.

[0064] The video processing library provides a series of functions and tools for loading, playing, processing, and analyzing video data. The anchor box mechanism is a technique used in object detection algorithms, especially in deep learning-based object detection models. Anchor boxes are a predefined set of rectangular boxes with different shapes and sizes, used to predict the positions of objects in an image. The receptive field refers to the size of the region of the input image that a neuron in the network can "see". Candidate positions usually refer to the image regions that may contain the target object.

[0065] Use a pre-set video device (such as a camera, video file, etc.) to obtain the input video stream. Use the video processing library (such as OpenCV) to extract images frame by frame from the video stream to obtain a series of image frames. Input the extracted image frames into a pre-set convolutional neural network and perform forward propagation. Each convolutional layer extracts features from the image frames to generate corresponding convolutional feature maps.

[0066] Based on the generated convolutional feature maps, use the region screening network to propose the bounding boxes and corresponding scores of candidate regions at each position through the anchor box mechanism. Sort the candidate regions in descending order according to the score values of each candidate region, and select the candidate regions in the top 10%. For each selected candidate region, determine its coordinate information and extract the corresponding image region in the image frame.

[0067] For each selected candidate region, based on the preset context information encoding module, determine the features of the surrounding regions of each candidate region by expanding the receptive field, so as to determine the context information of each candidate region. Combine the position information of the candidate region to determine the candidate position.

[0068] In this embodiment, by analyzing the video stream frame by frame and using the convolutional neural network to extract image features, video data can be processed and analyzed efficiently, and useful information for subsequent tasks can be extracted. Using the region screening network and the anchor box mechanism can accurately propose candidate regions, which helps to improve the accuracy of object detection and recognition. By analyzing the context information of each candidate region, the role and state of the target object in the scene can be better understood. In summary, this embodiment provides strong support for further image analysis and processing tasks by efficiently and accurately extracting useful information from the video stream and considering the context information, thus playing an important role in multiple application scenarios.

[0069] In an alternative embodiment,

[0070] Determining the contour of the target entity and performing segmentation through the fully connected layer and the fully convolutional layer, and obtaining the segmentation contour of the target entity includes:

[0071] Initialize the fully connected layer and the fully convolutional layer based on the candidate positions and the convolutional feature map, and add the convolutional feature map to the input interface of the fully connected layer;

[0072] The fully connected layer maps the convolutional feature map to a high-dimensional space through feature mapping and obtains the high-level semantic information corresponding to the convolutional feature map;

[0073] Add the obtained high-level semantic information to the fully convolutional layer, restore the high-level semantic information to the original resolution through upsampling operation, and generate the initial segmentation contour of the target entity through an activation function;

[0074] Fill and smooth the initial segmentation contour, compare the processed initial segmentation contour with the original image and evaluate the confidence. If the confidence is greater than the preset confidence threshold, add the initial segmentation contour as the segmentation contour of the target entity to the segmentation contour set and save it.

[0075] The fully connected layer is a layer in which each input node is fully connected to the output nodes. It is usually used in the last few layers of the network to integrate the features extracted by the previous convolutional layers and perform classification or regression analysis. The high-level semantic information refers to that as the network depth increases, the network can extract more abstract and complex features from the original input image, that is, high-level semantic features. Upsampling is a common image processing technique used to increase the resolution of an image. The initial segmentation contour refers to the edge contour of the target object preliminarily predicted by the model. The confidence threshold is the threshold that determines whether a prediction is accepted.

[0076] Based on the selected candidate positions and the generated convolutional feature map, initialize the fully connected layer and the fully convolutional layer, add the convolutional feature map to the input interface of the fully connected layer, and the fully connected layer maps the convolutional feature map to a high-dimensional space through feature mapping to obtain the high-level semantic information corresponding to the convolutional feature map;

[0077] Add the obtained high-level semantic information to the fully convolutional layer, restore the high-level semantic information to the original resolution through upsampling operation, and use an activation function to generate the initial segmentation contour of the target entity in the fully convolutional layer;

[0078] Fill and smooth the generated initial segmentation contour to improve the accuracy and quality of the contour, compare the processed segmentation contour with the original image, and evaluate the confidence of the contour. If the confidence is greater than the preset confidence threshold, confirm the segmentation contour as the final segmentation contour of the target entity and add it to the segmentation contour set for saving.

[0079] In this embodiment, by using a deep learning model to extract features and obtain high-level semantic information from an image, the content of the image can be understood more accurately, thereby improving the accuracy of target entity segmentation. By using a convolutional neural network to extract features from images of different sizes and shapes, it has good adaptability and can process various types of image data. By considering the context information of each candidate region, the relationship between the target entity and its surrounding environment can be better understood, which helps to improve the accuracy of segmentation. By evaluating the confidence of the segmentation result and comparing it with a preset threshold, the segmentation contour can be dynamically optimized to ensure the accuracy and reliability of the final result. In summary, this embodiment effectively improves the accuracy and automation of target entity segmentation by combining deep learning technology and image processing technology.

[0080] S2. Based on the segmentation contour, determine the foot position of the target entity in the bounding box through a cascaded pyramid network, calculate the vertical projection of the target entity based on the perspective projection principle to generate a target image. For the target image, perform human pose estimation through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human pose, map the two-dimensional human pose to the three-dimensional space to obtain a three-dimensional human pose, and perform time series analysis on the three-dimensional human pose to obtain action features.

[0081] The cascaded pyramid network is a neural network with a cascaded structure, usually used for target detection tasks. Through a series of cascaded detection modules, the detection of the bounding box is gradually refined. The bounding box is a rectangular box used to represent the position of the target in the image, usually determined by the coordinates of the upper left corner and the lower right corner. The perspective projection is a method of representing the projection of points in three-dimensional space onto a two-dimensional image. The vertical projection refers to projecting the information of an object or scene in the vertical direction onto a plane. The multi-scale processing mechanism refers to using features of different scales in image processing or computer vision tasks to enhance the robustness and generalization of the model. The action features refer to the features extracted in action recognition or analysis tasks to describe actions.

[0082] In an alternative embodiment,

[0083] The step of based on the segmentation contour, determining the foot position of the target entity in the bounding box through a cascaded pyramid network, and calculating the vertical projection of the target entity based on the perspective projection principle to generate a target image includes:

[0084] Extract the image features of the image frame based on the convolutional feature map, construct multiple convolutional layers and upsampling layers, generate multiple first feature maps with different resolutions by stacking convolutional kernels, and form a pyramid network. In the cascaded pyramid network, set multiple cascaded detection modules.

[0085] Based on the cascaded detection module, foot detection is performed on the first feature map in the order of increasing resolution. First, the approximate position of the feet of the target entity is obtained from the first feature map with a lower resolution, and the position of the feet is refined by detecting the first feature map with a higher resolution.

[0086] For each detected foot region, the bounding box of the detected foot region is adjusted through the bounding box regression technique until it completely coincides with the foot image of the target entity. Based on the position of the feet of the target entity, the position of the vertical projection point of the lowest point of the feet in the image on the ground is calculated to generate the target image.

[0087] The pyramid network is a neural network structure designed to process information at different scales or resolutions. The cascaded detection module is a technique used in a target detection system to improve accuracy by cascading multiple detectors, and each detector gradually refines the position information of the target.

[0088] A pyramid network is constructed using convolutional feature maps. Multiple feature maps with different resolutions are generated by stacking convolutional kernels and upsampling layers to form a pyramid network. The constructed pyramid network is used as the basis of the cascaded pyramid network. Multiple cascaded detection modules are set, and each module includes a foot detector and a corresponding bounding box regressor. The first cascaded detection module is used to detect the feet on the feature map with a lower resolution to obtain the approximate position of the feet of the target entity. For subsequent cascaded detection modules, the resolution of the feature map is gradually increased to refine the position of the feet.

[0089] For each detected foot region, the bounding box of the detected foot region is adjusted through the bounding box regression technique until it completely coincides with the foot image of the target entity. According to the position information of the foot region, the position of the vertical projection point of the lowest point of the feet in the image on the ground is calculated, and the target image is generated using the calculated position of the lowest point of the feet on the ground.

[0090] In summary, in this embodiment, the construction of the pyramid network allows the system to process images simultaneously at multiple resolutions, thereby better capturing the multi-scale features of the target entity, helping to cope with the changes of the target at different sizes and distances. The setting of the cascaded detection module enables the system to gradually refine the detection of the feet. Through multi-stage cascaded detection, the accuracy of the position of the target entity is improved. The bounding box regression technique finely adjusts the detected foot region, further improving the accuracy of the detection. By calculating the vertical projection point of the lowest point of the feet on the ground, the system can accurately determine the position of the target entity on the ground, thereby generating the target image. In summary, this embodiment realizes multi-scale foot detection, gradually refined positioning, improves the accuracy of the target position and the robustness of the system.

[0091] In an alternative embodiment,

[0092] For the target image, human pose estimation is performed through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human pose, the two-dimensional human pose is mapped to a three-dimensional space to obtain a three-dimensional human pose, and time series analysis is performed on the three-dimensional human pose to obtain action features including:

[0093] For the target image, based on a pre-set multi-scale processing mechanism, feature representations of the target image at multiple scales are extracted, and the features extracted at different scales are fused to obtain a fused feature map;

[0094] Based on the fused feature map, features are learned at different scales and images corresponding to target entities of different sizes are determined. Based on the images corresponding to the target entities, human pose estimation is performed to determine the two-dimensional human pose in the images corresponding to the target entities.

[0095] The internal and external parameters of the camera corresponding to the two-dimensional human pose are determined, the corresponding points of the contour in the two-dimensional human pose are normalized to obtain normalized coordinates, the camera projection model is applied to the normalized coordinates, the three-dimensional virtual coordinates corresponding to each normalized coordinate are calculated, and combined with the scale factor, the three-dimensional pose coordinates corresponding to the target entity are calculated to obtain the three-dimensional human pose;

[0096] The three-dimensional pose coordinates are tracked in consecutive frames, the temporal characteristics of the actions of the target entity are determined through a sliding window, and based on the temporal characteristics, the speed information and angle information corresponding to the actions are analyzed to determine the action features.

[0097] The two-dimensional human pose is used to describe the pose of the human body on a two-dimensional image plane, usually represented by the coordinates of key points, such as the positions of limb joints. The camera parameters include internal parameters (such as focal length, principal point coordinates) and external parameters (such as the rotation and translation matrices of the camera), which are used to map three-dimensional space points to a two-dimensional image. The camera projection model is a mathematical model used to describe how the camera projects three-dimensional space points onto a two-dimensional image. Common ones include perspective projection and orthographic projection. The scale factor is used to represent the proportional relationship between the actual size of an object in the image and the pixel coordinates, and can be calculated through the known size of the object and the corresponding pixel size.

[0098] Using a pre-set multi-scale processing mechanism, feature representations of the target image at low, medium, and high scales are extracted, the feature representations extracted at the three scales are fused to obtain a fused feature map. Based on the fused feature map, features are learned at low, medium, and high scales, images corresponding to target entities of different sizes are determined, and the determined images corresponding to the target entities are used for human pose estimation to obtain a two-dimensional human pose;

[0099] Based on the two-dimensional human pose, determine the internal and external parameters of the camera, normalize the corresponding points of the contour in the two-dimensional human pose to obtain the normalized coordinates, apply the camera projection model, calculate the three-dimensional virtual coordinates corresponding to each normalized coordinate, and combine with the scale factor to calculate the three-dimensional human pose corresponding to the target entity;

[0100] Track the three-dimensional pose coordinates in consecutive frames to form temporal information, process the temporal information through a sliding window to determine the action temporal characteristics of the target entity, and based on the temporal characteristics, analyze the speed information and angle information corresponding to the action to determine the action characteristics, such as jump height, rotation angle, etc., to form the final action characteristic description.

[0101] In this embodiment, through multi-scale feature extraction and fusion, the system can more comprehensively and comprehensively capture the feature information of the target entity at different scales, improve the global perception ability of the target, and based on the learned features and images at different scales, the system can accurately estimate the two-dimensional human pose of the target entity, providing a good basis for subsequent three-dimensional pose calculation. Using the internal and external parameters of the camera and the normalized coordinates, the system can accurately calculate the three-dimensional pose of the target entity, realizing an accurate mapping from the image space to the three-dimensional space. In summary, this embodiment realizes a comprehensive and accurate grasp of the pose and action of the target entity, providing a solid foundation for efficient human action analysis.

[0102] S3. Based on the action characteristics, combine with a preset action analysis model to construct a linear kernel function, and based on the linear kernel function, combine with motion action classification to determine the human activity of the target entity.

[0103] The linear kernel function is one of the commonly used kernel functions in the support vector machine (SVM). By performing a linear mapping in the input feature space, it maps the non-linear problem into a linear problem in a high-dimensional space, enabling the classifier in the high-dimensional space to more easily find a hyperplane for classification. The action analysis model is a computational model used to analyze and recognize human motion actions, and the motion action classification refers to classifying the input motion data and dividing it into predefined motion action categories.

[0104] In an alternative embodiment,

[0105] The determining the human activity of the target entity by combining with a preset action analysis model based on the action characteristics, constructing a linear kernel function based on the linear kernel function, and combining with motion action classification includes:

[0106] Based on the pre-acquired action characteristics and temporal characteristics, generate an empty set and add the known human actions to the empty set, and name the empty set activity information;

[0107] Based on the activity information, a linear kernel function corresponding to the action feature is constructed by combining a preset action analysis model;

[0108] Based on the linear kernel function, the action feature and the timing feature are mapped into a high-dimensional space to generate a high-dimensional action feature set. For each element in the high-dimensional action feature set, the Euclidean distance between each element is calculated by combining the action analysis model. Combining the position of the motion action classification in the current space, the action classification corresponding to each action feature is determined, and the human activity corresponding to the target entity is obtained based on the action classification.

[0109] Create an empty set named activity information, and add the pre-acquired known human action features to the activity information set;

[0110] Using the activity information set, combined with a preset action analysis model, a linear kernel function corresponding to the action feature is constructed. Using the constructed linear kernel function, the action feature and the timing feature are mapped into a high-dimensional space to generate a high-dimensional action feature set. For each element in the high-dimensional action feature set, the Euclidean distance between it and other elements is calculated by combining the action analysis model. Using a preset motion action classification model, the position of each action feature on the classification label is determined in the current high-dimensional space. Combining the motion action classification result, the human activity corresponding to the target entity is determined, and the human activity corresponding to the target entity determined by classification is output, and this activity is represented by a pre-defined action category label.

[0111] In this embodiment, by adding known action features to the activity information set, the system can efficiently utilize the pre-acquired action features to form a complete activity information library. Based on the linear kernel function, the system maps the action feature and the timing feature into a high-dimensional space, enabling the original features to be more comprehensively expressed in a more complex space. By calculating the Euclidean distance, the system can quantify the similarity between different action features, providing an effective basis for subsequent classification. Using a preset motion action classification model, the system accurately classifies the action features in the high-dimensional space, realizing the distinction of different actions. In summary, this embodiment can achieve the efficient processing and accurate classification of action features, providing strong support for the in-depth analysis of human activities.

[0112] In an alternative embodiment,

[0113] The method further includes training the action analysis model:

[0114] Collect real action data, randomly generate feature vectors according to the real action data, divide the randomly generated feature vectors into a training set and a test set, initialize the action analysis model and introduce a regularization parameter, and input the feature vectors in the training set into the action analysis model to generate a first decision boundary;

[0115] Based on the first decision boundary, determine the true positive rate and false positive rate corresponding to the action analysis model, obtain the classification probability corresponding to the action analysis model, based on the classification probability, combine the true class and the model prediction to determine the confusion matrix, based on the classification probability and the confusion matrix, dynamically update the action analysis model, add the test set to the updated action analysis model to obtain a second decision boundary, and based on the second decision boundary, repeat updating the feature vectors in combination with the regularization parameter until a preset stop count is reached.

[0116] Collect real action data from the actual scenario, including feature vectors and corresponding class labels. In the case of no real labels, randomly generate a part of the feature vectors for subsequent training and testing, and divide the real action data and the randomly generated feature vectors into a training set and a test set;

[0117] Use the feature vectors and real labels of the training set to initialize the action analysis model and introduce a regularization parameter. Input the feature vectors of the training set into the initialized action analysis model to generate a first decision boundary. Based on the first decision boundary, determine the true positive rate, false positive rate and classification probability of the action analysis model, combine the true class and the model prediction, calculate the confusion matrix, and dynamically update the action analysis model according to the classification probability and the confusion matrix;

[0118] Add the test set to the updated action analysis model. Based on the updated model, generate a second decision boundary. Combine the second decision boundary and the regularization parameter to repeat updating the randomly generated feature vectors until a preset stop count is reached. Check whether the preset stop count is reached. If so, end the training process.

[0119] In this embodiment, by collecting real action data, the system can obtain the action information in the real scenario, improve the adaptability of the model in the actual environment. For the real action data, the system randomly generates feature vectors and divides them into a training set and a test set, ensuring the diversity of the data and the generalization ability of the model. According to the confusion matrix and the classification probability, the system generates a second decision boundary, improving the ability of the model to distinguish action features. In summary, this embodiment realizes the dynamic adjustment and optimization of the action analysis model, thereby improving the accuracy and robustness of the model in the real scenario.

[0120] Figure 2This is a schematic structural diagram of the personnel activity analysis system based on image recognition according to an embodiment of the present invention. As Figure 2 shown, the system includes:

[0121] A first unit, configured to obtain a video stream, extract image frames from the video stream, extract features in the image frames through a convolutional network and generate a convolutional feature map, determine candidate regions of a target entity through a region screening network for the convolutional feature map, combine context information corresponding to the candidate regions to obtain candidate positions, and determine the contour of the target entity and perform segmentation through a fully connected layer and a fully convolutional layer to obtain a segmentation contour of the target entity;

[0122] A second unit, configured to, based on the segmentation contour, determine the foot position of the target entity in a bounding box through a cascaded pyramid network, calculate the vertical projection of the target entity based on the principle of perspective projection to generate a target image, perform human pose estimation on the target image through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human pose, map the two-dimensional human pose to a three-dimensional space to obtain a three-dimensional human pose, and perform time series analysis on the three-dimensional human pose to obtain action features;

[0123] A third unit, configured to, based on the action features, combine a preset action analysis model to construct a linear kernel function, and based on the linear kernel function, combine motion action classification to determine the human activity of the target entity.

[0124] In a third aspect of the embodiments of the present invention,

[0125] There is provided an electronic device, including:

[0126] A processor;

[0127] A memory for storing instructions executable by the processor;

[0128] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0129] In a fourth aspect of the embodiments of the present invention,

[0130] There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0131] The present invention can be a method, a device, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0132] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A personnel activity analysis method based on image recognition, characterized in that: include: Acquire a video stream, extract image frames from the video stream, extract features from the image frames through a convolutional network and generate a convolutional feature map, determine a candidate region of a target entity through a region screening network using the convolutional feature map, obtain a candidate position by combining context information corresponding to the candidate region, determine the contour of the target entity through a fully connected layer and a fully convolutional layer, and perform segmentation to obtain a segmented contour of the target entity; Based on the segmented contour, the foot position of the target entity in the bounding box is determined through a cascade pyramid network, the vertical projection of the target entity is calculated based on the perspective projection principle to generate a target image, for the target image, human body posture estimation is performed through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human body posture, the two-dimensional human body posture is mapped to a three-dimensional space to obtain a three-dimensional human body posture, and the three-dimensional human body posture is analyzed in time series to obtain action features; Based on the motion features and in combination with a preset motion analysis model, a linear kernel function is constructed, and based on the linear kernel function and in combination with motion classification, the human body activity of the target entity is determined; Extracting image features of the image frame based on the convolution feature map, constructing multiple convolution layers and upsampling layers, generating multiple first feature maps of different resolutions by stacking convolution kernels, and forming a pyramid network, in which multiple cascade detection modules are set; Based on the cascade detection module, the first feature map is used to detect the foot in an ascending order of resolution, and the approximate position of the foot of the target entity is first obtained by using the first feature map with a lower resolution, and the position of the foot is refined by detecting the first feature map with a higher resolution; For each detected foot area, the bounding box of the detected foot area is adjusted by bounding box regression technology until it completely overlaps with the foot image of the target entity. Based on the foot position of the target entity, the position of the vertical projection point of the lowest point of the foot in the image on the ground is calculated to generate the target image.

2. The method according to claim 1, characterized in that The step of acquiring a video stream, extracting an image frame from the video stream, extracting features from the image frame through a convolutional network and generating a convolutional feature map, determining a candidate region of a target entity through a region screening network, and combining context information corresponding to the candidate region to obtain a candidate position includes: Obtain an input video stream through a preset video device, extract images frame by frame in the video stream through a video processing library to obtain multiple image frames, input the image frames into a preset convolutional neural network, forward propagate the image frames in the convolutional neural network, and extract features from the image frames in each convolutional layer to generate a corresponding convolutional feature map; Based on the convolutional feature map, the region screening network proposes the border and corresponding score of the candidate region for each position through the anchor frame mechanism, sorts the candidate regions in descending order according to the score of each candidate region, selects the candidate regions in the top 10%, determines the coordinate information for each selected candidate region, and extracts the corresponding image region in the image frame; For each selected candidate region, based on a preset context information encoding module, the features of the surrounding area of ​​each candidate region are determined by expanding the receptive field, the context information of each candidate region is determined, and the candidate position is determined in combination with the position information of the candidate region.

3. The method according to claim 1, characterized in that Determining the contour of the target entity and segmenting it through the fully connected layer and the fully convolutional layer to obtain the segmented contour of the target entity includes: Initializing the fully connected layer and the fully convolutional layer based on the candidate position and the convolutional feature map, and adding the convolutional feature map to the input interface of the fully connected layer; The fully connected layer maps the convolution feature map to a high-dimensional space through feature mapping and obtains high-level semantic information corresponding to the convolution feature map; Adding the acquired high-level semantic information to the full convolutional layer, restoring the high-level semantic information to the original resolution through an upsampling operation, and generating an initial segmentation outline of the target entity through an activation function; The initial segmentation contour is filled and smoothed, the processed initial segmentation contour is compared with the original image and the confidence is evaluated. If the confidence is greater than a preset confidence threshold, the initial segmentation contour is added to the segmentation contour set as the segmentation contour of the target entity and saved.

4. The method according to claim 1, characterized in that: For the target image, human body posture estimation is performed through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human body posture, the two-dimensional human body posture is mapped to a three-dimensional space to obtain a three-dimensional human body posture, and a time series analysis is performed on the three-dimensional human body posture to obtain action features including: For the target image, based on a pre-set multi-scale processing mechanism, feature representations of the target image at multiple scales are extracted, and features extracted at different scales are fused to obtain a fused feature map; Based on the fused feature map, features are learned at different scales and images corresponding to target entities of different sizes are determined, human posture estimation is performed based on the images corresponding to the target entities, and a two-dimensional human posture in the images corresponding to the target entities is determined, Determine the internal and external parameters of the camera corresponding to the two-dimensional human body posture, normalize the corresponding points of the contour in the two-dimensional human body posture to obtain normalized coordinates, apply the camera projection model to the normalized coordinates, calculate the three-dimensional virtual coordinates corresponding to each normalized coordinate, and calculate the three-dimensional posture coordinates corresponding to the target entity in combination with the scale factor to obtain the three-dimensional human body posture; The three-dimensional posture coordinates are tracked in continuous frames, and the timing characteristics of the action of the target entity are determined through a sliding window. Based on the timing characteristics, speed information and angle information corresponding to the action are analyzed to determine the action characteristics.

5. The method according to claim 1, characterized in that: The step of constructing a linear kernel function based on the motion feature and in combination with a preset motion analysis model, and determining the human activity of the target entity based on the linear kernel function and in combination with motion classification includes: Based on the pre-acquired action features and time series features, an empty set is generated and known human actions are added to the empty set, and the empty set is named activity information; Based on the activity information, a linear kernel function corresponding to the action feature is constructed in combination with a preset action analysis model; Based on the linear kernel function, the motion features and the timing features are mapped to a high-dimensional space to generate a high-dimensional motion feature set. For each element in the high-dimensional motion feature combination, combined with the motion analysis model, the Euclidean distance between each element is calculated. Combined with the position of the motion action classification in the current space, the action classification corresponding to each action feature is determined, and the human body activity corresponding to the target entity is obtained based on the action classification.

6. The method according to claim 1, characterized in that The method further comprises training the motion analysis model: Collecting real action data, randomly generating feature vectors according to the real action data, dividing the randomly generated feature vectors into a training set and a test set, initializing the action analysis model and introducing a regularization parameter, inputting the feature vectors in the training set into the action analysis model, and generating a first decision boundary; Based on the first decision boundary, determine the true positive rate and false positive rate corresponding to the action analysis model to obtain the classification probability corresponding to the action analysis model; based on the classification probability, determine the confusion matrix in combination with the true category and the model prediction; based on the classification probability and the confusion matrix, dynamically update the action analysis model, add the test set to the updated action analysis model to obtain the second decision boundary; based on the second decision boundary, repeatedly update the feature vector in combination with the regularization parameter until a preset number of stop times is reached.

7. A personnel activity analysis system based on image recognition, used to implement the personnel activity analysis method based on image recognition as described in any one of claims 1 to 6, characterized in that: include: The first unit is used to obtain a video stream, extract an image frame in the video stream, extract features in the image frame through a convolutional network and generate a convolutional feature map, determine a candidate region of a target entity through a region screening network, combine context information corresponding to the candidate region to obtain a candidate position, determine the contour of the target entity through a fully connected layer and a fully convolutional layer and perform segmentation to obtain a segmented contour of the target entity; The second unit is used to determine the foot position of the target entity in the bounding box through a cascade pyramid network based on the segmented contour, calculate the vertical projection of the target entity based on the perspective projection principle, generate a target image, estimate the human body posture of the target image through a pre-introduced multi-scale processing mechanism to obtain a two-dimensional human body posture, map the two-dimensional human body posture to a three-dimensional space to obtain a three-dimensional human body posture, and perform time series analysis on the three-dimensional human body posture to obtain action features; The third unit is used to construct a linear kernel function based on the action feature in combination with a preset action analysis model, and determine the human activity of the target entity based on the linear kernel function in combination with motion classification.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.