A method, system, computer device and storage medium for detecting motion postures
Through the rectangular expansion deformable convolution mechanism and adaptive feedforward neural network, combined with multiple receptive field feature maps and center point detectors, the false detection and missed detection problems caused by pedestrian pose changes are solved, and high-precision pedestrian motion posture detection is achieved.
Patent Information
- Application Number
- CN202510660071.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-22
AI Technical Summary
When existing pedestrian detection technology deals with pedestrians with different postures, there are problems of mis-detection and missed detection, making it difficult to achieve high-precision motion posture detection.
The rectangular expansion deformable convolution mechanism and an adaptive feedforward neural network are used to process the rectangular sampling grid of multiple branches in parallel and the learnable offset function, combining the linear layer to capture local details and global information, generate multiple receptive field feature maps, and use a central point-based detector to perform pedestrian feature detection.
It improves the accuracy of pedestrian motion posture detection, reduces false detection and missed detection, and enhances the detection ability of pedestrians with different postures.
Smart Images

Figure CN120183049B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method, a system, a computer device and a storage medium for detecting a motion posture. Background Art
[0002] Pedestrian detection is an important research direction in the field of computer vision, which aims to identify and locate all pedestrians from images or videos. This technology plays an important role in many application scenarios, such as intelligent monitoring, autonomous driving vehicles, human-computer interaction interfaces, etc.
[0003] The goal of pedestrian detection is to accurately identify the presence of pedestrians from the input visual data (such as pictures or video streams captured by cameras) and determine their positions (usually in the form of bounding boxes). This involves complex algorithms and technologies to handle various challenges, including but not limited to the pose changes of pedestrians, occlusion, cluttered backgrounds, and changes in lighting conditions.
[0004] With the wide application of pedestrian detection in real-world scenarios such as intelligent monitoring, human-computer interaction, and autonomous driving, the requirement for detection accuracy is getting higher and higher. However, in existing pedestrian detection technologies, the appearance of pedestrians with different postures varies greatly, and the problem of pedestrian pose changes often leads to misdetection and missed detection by detectors. To accurately capture pedestrian targets in images or videos, there is an urgent need for a motion posture detection method with higher detection accuracy. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a method for detecting a motion posture, aiming to improve the accuracy of detecting the motion postures of pedestrians and avoid misdetection and missed detection caused by pedestrian pose changes.
[0006] The embodiments of the present invention are implemented as follows. A method for detecting a motion posture, the method for detecting a motion posture includes:
[0007] Obtain image data;
[0008] Divide the image data into several parallel branches, and use different rectangular dilation rates in each of the branches to generate rectangular sampling grids, and combine a learnable offset function to achieve adaptive offset of sampling points, capturing local detail information of the branches; use a linear layer to capture the global information of each branch; cascade the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map;
[0009] Detect pedestrian features in the multi-receptive field feature map.
[0010] Further, the specific method for obtaining the image data is as follows: Obtain a number of feature maps with different resolutions, perform unified sampling using the bilinear interpolation algorithm to obtain a common resolution, and concatenate them in the channel dimension to obtain the image data.
[0011] Further, different numbers of hole elements are inserted between the sampling points in the convolution kernel of the rectangular sampling grid to form a rectangular sampling range for guiding the network to capture pedestrian features in the shape of a rectangle. The pedestrian features are specifically:
[0012] ;
[0013] Among them, and respectively represent the vertical and horizontal dilation rates of the i-th branch.
[0014] Further, an adaptive offset of the sampling points is achieved using a learnable offset function to obtain the output feature map of the i-th branch:
[0015] ;
[0016] Among them, represents the output feature map of the i-th branch, represents the position of the pixel point in the entire image data, is the predefined position of the n-th point, represents the weight of the convolution kernel, represents the learnable offset, represents the input feature map obtained by the i-th branch through the rectangular dilation deformable convolution mechanism.
[0017] Further, the specific operation process for capturing the global information of each branch using the linear layer is as follows:
[0018] ;
[0019] Among them, represents the output feature map of the i-th branch, BN(·) represents the batch normalization operation, ReLU(·) represents the activation function, represents the first linear layer, represents the second linear layer.
[0020] Further, the output feature maps of all branches are concatenated in the channel dimension. The specific operation is as follows:
[0021] ;
[0022] Among them, represents the multi-receptive field feature map; in each branch, represents the feature map obtained by performing rectangle dilation deformable convolution and adaptive offset processing on the i-th branch; represents the feature map obtained by processing through an adaptive feed-forward neural network; represents the output feature map of each branch; represents the input feature map; represents a two-dimensional convolutional layer with a convolution kernel of ; represents concatenating the feature maps in the channel dimension.
[0023] Furthermore, to detect pedestrian features in the multi-receptive field feature map, a center point-based detector is used, and the steps are as follows:
[0024] Use the multi-receptive field feature map after dimensionality reduction processing as the processing target;
[0025] Predict the target center point, target scale, and target center point offset respectively. Generate a target candidate box for the input image according to the predicted target center and target scale, and fine-tune the position of the target center point through the target center point offset;
[0026] Construct a ground truth label for each prediction branch in the detector, and apply a two-dimensional Gaussian mask G(∙) at the position of each positive sample, which is expressed as follows:
[0027] ;
[0028] where (i,j) represents the position of the center point output by the network, K represents the number of targets in the input image, are the coordinates, width, and height of the k-th target center point, and the variances and of the two-dimensional Gaussian distribution are proportional to the width and height of the target respectively;
[0029] Define the target scale as the width and height of the target. The position of the k-th positive sample is assigned to the of the k-th target, and is also assigned to all negative sample points within a radius of 2 from the positive sample point. Other positions are all assigned 0;
[0030] The position of the center point of the input image is mapped to the position in the output image, and the ground truth of the target center point offset is defined as:
[0031] ;
[0032] where denotes the floor operation, that is, taking the largest integer not greater than this value, and r is the downsampling factor.
[0033] Another object of the embodiments of the present invention is a motion posture detection system, characterized in that the motion posture detection system executes the motion posture detection method, and the motion posture detection system includes:
[0034] A backbone network to obtain image data;
[0035] An adaptive feature enhancement module that divides the image data into several parallel branches. Each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, and combines a learnable offset function to achieve adaptive offset of sampling points, capturing local detail information of the branch; uses a linear layer to capture the global information of each branch; cascades the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map;
[0036] A detector to detect pedestrian features in the multi-receptive field feature map.
[0037] Another object of the embodiments of the present invention is a computer device, characterized in that it includes a memory and a processor. A computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the motion posture detection method.
[0038] Another object of the embodiments of the present invention is a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the processor executes the steps of the motion posture detection method.
[0039] A motion posture detection method provided by the embodiments of the present invention first proposes a rectangular dilation deformable convolution mechanism, which combines the advantages of the wide receptive field range of dilation convolution and the adaptive sampling ability of deformable convolution. To avoid the limitation of the local receptive field range of the method of the present invention and capture global features, an adaptive feedforward neural network is designed. Based on the previous step, the global information is modeled through a linear layer, significantly expanding the feature representation ability. On this basis, the present invention further proposes a star mechanism, which divides the image data into several independent branches to achieve the sequential cascade of the rectangular dilation deformable convolution mechanism module and the adaptive feedforward neural network within each branch, and at the same time makes full use of the local detail capture ability and the global information modeling ability, enabling the method to accurately and parallelly process targets of different scales.
[0040] The present invention is mainly embodied in several aspects: First, different numbers of hole elements are inserted between sampling points to design an innovative rectangular sampling grid scheme to guide the network to pay more attention to rectangular pedestrian targets; Second, on the basis of the rectangular sampling grid, the deformable convolution is further combined to innovatively propose a rectangular dilated deformable convolution (RDDC) mechanism to achieve adaptive sampling within the rectangular receptive field to cope with the complex pose changes of pedestrians; Third, the Stars mechanism is designed to construct the AFE-Module, and multi-scale feature maps are processed through a four-branch parallel structure to finally obtain multi-receptive field feature maps. Description of the Drawings
[0041] Figure 1 It is an application environment diagram of the motion pose detection method provided by an embodiment of the present invention;
[0042] Figure 2 It is a pedestrian detection network (MAFE-Net) based on multi-receptive field adaptive feature enhancement provided by an embodiment of the present invention;
[0043] Figure 3 It is a structural diagram of an adaptive feature enhancement module (AFE-Module) provided by an embodiment of the present invention;
[0044] Figure 4 It is an example diagram of a rectangular dilated deformable convolution module (RDDC-Block) provided by an embodiment of the present invention;
[0045] Figure 5 It is a schematic diagram showing the changes in the receptive field range and sampling point positions of four sub-branches and their fused feature maps provided by an embodiment of the present invention;
[0046] Figure 6 It is an internal structure block diagram of a computer device in one embodiment. Detailed Embodiments
[0047] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0048] Figure 1 An application environment diagram of a motion pose detection method is provided by an embodiment of the present invention. As Figure 1 shown, in this application environment, an image acquisition device and a computer device are included.
[0049] An image acquisition device refers to a hardware tool for capturing, recording or generating digital images or videos, which can be a camera, a scanner, a digital camera or an industrial camera, etc.
[0050] The computer device can be a smartphone, a tablet computer, a laptop computer, or a desktop computer, or it can be a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, cloud storage, and CDN. The image acquisition device and the computer device can be connected through a network, which is not limited in the present invention.
[0051] As Figure 2 and Figure 3 shown, in the first embodiment, a motion posture detection method is proposed. In this embodiment, the method is mainly illustrated by applying it to the computer device in the above Figure 1 . A motion posture detection method may specifically include the following steps:
[0052] Step S202, obtain image data.
[0053] In this embodiment, as Figure 2 shown, the required image data is obtained in the backbone network (HRNet). The image acquisition device inputs a series of images into the backbone network, and the image data can be obtained only after being processed by the backbone network. In this embodiment, the image acquisition device inputs a 3-channel image with a size of H×W, and the image is subjected to feature extraction through the backbone network. The backbone network is a parallel structure, consisting of four branches with different resolutions. By parallel processing multiple resolution branches, the backbone network can combine the powerful semantic information of the feature map with accurate position information, promoting information interaction between different branches. The backbone network finally outputs four feature maps, and the pixel ratios are 4 times, 8 times, 16 times, and 32 times the input image ratio respectively. The feature map with a high downsampling rate contains rich deep semantic information, but due to its low resolution and relatively lack of details, it is not ideal for detection. In addition, considering that small-scale pedestrian targets account for a relatively large proportion in the pedestrian detection task, in this embodiment, the bilinear interpolation algorithm is used to uniformly upsample these four output feature maps to a common resolution and connect them in the channel dimension to fully integrate multi-scale information.
[0054] Step S204, divide the image data into several parallel branches, and each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, and combines a learnable offset function to achieve adaptive offset of the sampling points, capturing the local detail information of the branch; use a linear layer to capture the global information of each branch; cascade the output feature maps of all the branches in the channel dimension to obtain a multi-receptive field feature map.
[0055] In this embodiment, as Figure 2As shown, the method adopted is a pedestrian detection network (MAFE-Net) based on multi-receptive field adaptive feature enhancement, aiming to enhance pedestrian features in the feature map from the perspective of feature extraction and make the network pay more attention to pedestrian targets. Specifically, as Figure 3 shown, first, a rectangular dilated deformable convolution (RDDC) mechanism is proposed. Based on this, a rectangular dilated deformable convolution module (RDDC-Block) is constructed, which combines the advantages of the wide receptive field range of dilated convolution and the adaptive sampling ability of deformable convolution. RDDC first uses the idea of inserting hole elements between sampling points by dilated convolution and designs a unique rectangular sampling grid scheme based on the common aspect ratios of pedestrians, thus expanding the sampling range to be closer to the actual bounding box of pedestrians. Then, based on this sampling grid scheme, an adaptive offset function of deformable convolution is introduced, enabling the sampling points to be dynamically adjusted according to the input feature map of different pedestrian postures. However, relying solely on the RDDC-Block is still limited by the local receptive field range and cannot fully capture global features. Therefore, the present invention designs an adaptive feedforward neural network (AFFN-Block), which models global information through linear layers and significantly expands the feature representation ability of the RDDC-Block. On this basis, the present embodiment proposes a star-shaped (Stars) mechanism and constructs an adaptive feature enhancement module (AFE-Module). The Stars mechanism uses four independent sub-branches to achieve the sequential cascade of the RDDC-Block and AFFN-Block within each branch, and at the same time makes full use of the local detail capture ability of the RDDC-Block and the global information modeling ability of the AFFN-Block, enabling the AFE-Module to process targets of different scales in parallel.
[0056] Step S206: Detect pedestrian features in the multi-receptive field feature map.
[0057] In this embodiment, the multi-receptive field feature map is detected in the detector. MAFE-Net adopts a center-based detector. After the feature map containing target fine-grained information enters the detector, it first undergoes a 3×3 Conv layer for dimensionality reduction, reducing the channel dimension to 256. Then it passes through three parallel 1×1 Conv layers, predicting the target center point (center), target scale (scale), and target center point offset (offset) respectively. Finally, target candidate boxes of the input image are generated based on the predicted target center and target scale, and the center point position is fine-tuned through the target center point offset information, thereby further improving the performance of the detector.
[0058] In an embodiment, it is necessary to construct ground truth labels for each prediction branch in the detector. For the ground truth label of the target center point, since it is difficult to determine the accurate center point position of a target, in order to reduce the uncertainty of a large number of negative samples around the positive samples, a two-dimensional Gaussian mask G(∙) is applied at the position of each positive sample, and the formula is expressed as follows:
[0059] ;
[0060] where (i,j) represents the position of the center point output by the network, K represents the number of targets in the input image, are the coordinates, width, and height of the center point of the k-th target, and the variances and of the two-dimensional Gaussian distribution are respectively proportional to the width and height of the target;
[0061] The target scale is defined as the width and height of the target. For its ground truth label, the position of the k-th positive sample is assigned to the of the k-th target, and in order to reduce the error caused by point prediction, it is also assigned to all negative sample points within a radius of 2 from the positive sample point, and other positions are all assigned 0;
[0062] The center point position of the input image is mapped to the position in the output image, and the ground truth label of the target center point offset is defined as:
[0063] ;
[0064] where, represents the floor operation, that is, taking the largest integer not greater than this value, and r is the downsampling factor.
[0065] In the second embodiment, step S202 specifically includes steps S302 to S308. As Figure 3 shown, steps S302 to S308 of this embodiment are executed by the AFE-Module, and the AFE-Module is composed of the RDDC-Block and the AFFN-Block; the RDDC-Block executes steps S302 and S304, the AFFN-Block executes step S306, and step S308 is about the Stars mechanism of this method.
[0066] Step S302, different numbers of hole elements are inserted between the sampling points in the convolution kernel of the rectangular sampling grid to form a rectangular sampling range for guiding the network to capture the pedestrian features in the rectangular form, and the pedestrian features are specifically:
[0067] ;
[0068] ;
[0069] Among them, and respectively represent the vertical and horizontal dilation rates of the \(i\)-th branch. It should be noted that when and are both set to 1, they correspond to the standard grid \(R(1, 1)\) of the standard convolution. As Figure 2 shown, different dilation rates and are adopted for the four sub-branches (\(i = 1, 2, 3, 4\)) in the AFE-Module, and the dilation rates are \((8, 4)\), \((6, 3)\), \((4, 2)\), and \((2, 1)\) respectively. These dilation rates are called rectangular dilation rates (RDR), which are selected to achieve a fixed ratio of 2:1 dilation in the vertical and horizontal dimensions.
[0070] In this embodiment, in order to reduce the number of parameters and computational cost of the entire module, first, a \(1\times1\) convolutional layer is applied to the output feature map from the backbone network , so as to obtain the input feature map of the RDDC-Block:
[0071] ;
[0072] Pedestrians appear in the image. Regardless of their size or resolution, their predicted bounding boxes always appear as rectangles. Since the aspect ratio of most detection boxes is approximately 2:1, this method proposes a rectangular dilation deformable convolution (RDDC) mechanism to extract pedestrian features, which improves the flexibility of the network in processing pedestrian postures. The implementation of RDDC is as Figure 3 shown. Constrained by the receptive field range of the standard convolution and the inflexibility of sampling points, first, a rectangular sampling grid is generated using the rectangular dilation rate (RDR) to achieve the expansion of the rectangular receptive field of the traditional convolution operator. Then, a learnable offset function is introduced to achieve the adaptive offset of the sampling points. The advantage of this mechanism is that it effectively reduces the interference of background information and enables the network to focus on the features of the pedestrian target.
[0073] Step S304, use the learnable offset function to achieve the adaptive offset of the sampling points, and obtain the output feature map of the \(i\)-th branch:
[0074] ;
[0075] Among them, represents the output feature map of the \(i\)-th branch, Indicates the position of the pixel point in the entire image data, is the predefined position of the nth point, represents the weight of the convolutional kernel, represents the learnable offset, represents the input feature map obtained by the i-th branch through the rectangular dilation deformable convolution mechanism.
[0076] In this embodiment, a learnable deformable convolution sampling point offset is introduced , which can flexibly adjust the sampling point position. For the learnable sampling point offset in the 3×3 convolution operation, each sampling point has an offset value in both the vertical and horizontal directions. Since there are a total of 9 sampling points, this will generate 18 offset values in the vertical and horizontal directions. To obtain the learnable offset , a 3×3 standard convolution with 18 output channels is specifically designed, as shown in Figure 2 the RDDC-Block. This convolution effectively obtains these offsets, thereby realizing the adaptive adjustment of the sampling points. Finally, the weight is weighted and summed with the pixel values corresponding to the offset sampling points to obtain the output feature map of the i-th branch. This innovative design not only effectively solves the sparse sampling problem brought by the dilated convolution, but also improves the flexibility of the convolution sampling.
[0077] Step S306, using the linear layer to capture the global information of each branch, the specific operation process is as follows:
[0078] ;
[0079] Among them, represents the output feature map of the i-th branch, BN(·) represents the batch normalization operation, ReLU(·) represents the activation function, represents the first linear layer, represents the second linear layer.
[0080] Although the RDDC-Block effectively expands the receptive field, it is still limited in modeling long-distance global information. To solve this problem, this embodiment designs a new AFFN-Block, which consists of a linear layer and an activation function and is specifically used to capture global information. Through this design, the network can not only capture local fine-grained information, but also establish the dependency relationship between global information.
[0081] As shown in the AFFN-Block in Figure 3 , first, for the input Execute the BatchNorm (BN, Batch Normalization) and ReLU functions. Then, apply the first linear layer to the feature map , and the number of output neurons is four times the number of input neurons, which helps the network learn more complex feature representations. Subsequently, the ReLU activation function is adopted to introduce non-linear features into the network. Finally, the number of neurons is restored to the number of input neurons of through the second linear layer, ensuring the effective integration of information. Due to the information interaction between neurons in the linear layer, the AFFN-Block can integrate the information in the entire spatial dimension of the input features, thereby capturing global information.
[0082] Step S308, concatenate the output feature maps of all the branches in the channel dimension, and the specific operations are as follows:
[0083] ;
[0084] Among them, represents the multi-receptive field feature map; in each of the branches, represents the feature map of the i-th branch after rectangular dilated deformable convolution and adaptive offset processing, represents the feature map processed by the adaptive feed-forward neural network, represents the output feature map of each branch, represents the input feature map, represents a two-dimensional convolutional layer with a convolution kernel of , represents the operation of concatenating feature maps in the channel dimension.
[0085] In this embodiment, the Stars mechanism can further capture the multi-receptive fields of semantic information. As shown in the AFE-Module of Figure 3 , first construct a four-branch parallel structure, where the RDDC-Block and AFFN-Block are sequentially embedded in each branch. Different RDRs are applied to each branch to expand the range of the receptive field and enrich the feature diversity. Specifically, the values of RDRs used in each branch are: (8, 4), (6, 3), (4, 2), and (2, 1). Then connect all the branches in the channel dimension to achieve effective multi-scale information fusion and obtain the multi-receptive field feature map. Through the above multi-scale design, although the aspect ratios of some pedestrian detection boxes are not 2:1, the AFE-Module can still process them by fusing complementary information from different sub-branches.
[0086] Figure 4It is an example diagram of an RDDC-Block when the RDR is (8, 4). Figure 4 In (a) of Figure 4 , the sampling point positions and receptive field ranges of a 3×3 standard convolution are shown; Figure 4 In (b) of Figure 4 , the expansion process of the receptive field is shown; Figure 4 In (c) of Figure 4 , the offset process of the sampling points is shown; where Figure 4 The boxed areas in (b) and (c) of Figure 4 represent the expanded rectangular receptive field ranges, Figure 4 The matrix in (b) of Figure 4 represents the sampling grid . Figure 5 It represents a schematic diagram of the changes in the receptive field ranges and sampling point positions of the four sub-branches and their fused feature maps in the Stars mechanism. Figure 5 In (a) to (d) of Figure 5 , the four sub-branches of the RDDC-Block with the adopted RDRs of (8, 4), (6, 3), (4, 2), and (2, 1) are respectively represented; Figure 5 In (e) of Figure 5 , the change in the multiple receptive field ranges of the fused feature map is represented. Figure 5 It shows the changes in the sampling point positions and receptive field ranges of each sub-branch. This design enhances the feature representation while ensuring the integrity and accuracy of pedestrian information with different scales and poses.
[0087] Finally, the output feature maps of all branches are concatenated in the channel dimension to form a feature map with multiple receptive fields and rich in deep semantic information. Then, by using a 1×1 convolutional layer to adjust the number of channels and perform a residual connection with the input feature map to obtain the final output feature map . This design not only speeds up the convergence rate of the network but also improves the generalization of the AFE-Module.
[0088] In the third embodiment, the loss function of this method is calculated based on the first and second embodiments.
[0089] The total loss function in this application consists of three parts: the target center point loss , the target scale loss , and the target center point offset loss . Since the number of target center points in the input image is less than that of non-target center points, resulting in an imbalance between positive and negative samples, which is not conducive to the training of the network model. Therefore, Focal Loss is used to solve the problem of imbalance between positive and negative samples. For this reason, the center point classification loss function is defined as:
[0090] ;
[0091] Among them, K represents the number of targets in an image, and W and H respectively represent the width and height of the image. represents the probability that the predicted coordinate point (i, j) belongs to the center point of the target. β and γ are two hyperparameters, which are respectively set as β = 4 and γ = 2 in this application. represents the Gaussian heat map of the coordinate point (i, j).
[0092] The SmoothL1 loss function is used to calculate the target scale loss and the target center point offset loss, and its formula is:
[0093] ;
[0094] Among them, respectively represent the predicted value and the true value of the k-th target scale. respectively represent the predicted value and the true value of the k-th target center point offset.
[0095] To sum up, the total loss function L is defined as:
[0096] ;
[0097] Among them, respectively represent the weights of the target center point loss, the scale loss, and the center point offset loss. In this application, they are respectively set as 0.01, 1, and 0.1.
[0098] In the fourth embodiment, based on the methods of the above three embodiments, specific experiments and results are given.
[0099] The MAFE-Net algorithm proposed in this embodiment is implemented using PyTorch and runs on four A100 PCIE-40GB-GPU devices. The backbone network HRNet of this algorithm uses the model weights pre-trained on the ImageNet dataset. In addition, for the object detection task mentioned in this application, the Adam optimization algorithm is adopted. On the CityPersons dataset for pedestrian detection in object detection, the size of the input image is set to 640×1280, the number of iterations in the training stage is set to 150, the batch size is set to 16, and the initial learning rate is set to , and the learning rate is multiplied by 0.1 every 50 iterations. This embodiment uses the average miss rate as the evaluation index.
[0100] The evaluation subsets in CityPersons are divided into low-occlusion and multi-scale subsets to more fairly compare the effectiveness of methods in dealing with multi-pose pedestrians. On the one hand, for targets with pixel values greater than 50, those with target visibility rates in the range of [0.65, 1] are classified under the reasonable occlusion subset (Reasonable), those in the range of [0.65, 0.9] are classified under the partial occlusion subset (Partial), and those in the range of [0.9, 1] are classified under the bare occlusion subset (Bare). On the other hand, for targets with visibility rates greater than 0.65, those with pixel value ranges in [50, 75] are classified as the small-scale subset (Small), those in [75, 100] are classified as the medium-scale subset (Medium), and those with ranges greater than 100 are classified as the large-scale subset (Large).
[0101] Table 1 presents the experimental results of this embodiment on the CityPersons dataset and compares the average miss detection rates with existing detection algorithms in dealing with multi-pose pedestrians. The average miss detection rate of the MAFE-Net algorithm on all low-occlusion and multi-scale subsets is lower than that of other algorithms. In particular, the average miss detection rates on the reasonable occlusion subset (Reasonable) and the small-scale subset (Small) are 9.0% and 10.3% respectively, achieving improvements of 1.4% and 3.9% compared to the CSP algorithm using the same backbone network (HRNet). Even when compared with the PEN method also used to deal with pedestrian pose problems, MAFE-Net also shows strong competitiveness. The experimental results fully demonstrate the effectiveness of the adaptive feature enhancement module in MAFE-Net for multi-pose pedestrian detection.
[0102] Table 1 Comparison of the average miss detection rates between MAFE-Net and existing methods on CityPersons
[0103]
[0104] In the fifth embodiment, a motion pose detection system is provided, characterized in that the motion pose detection system includes the steps and sub-steps of the motion pose detection method described in the first to third embodiments of the present invention, and the motion pose detection system includes:
[0105] A backbone network for acquiring image data;
[0106] Adaptive feature enhancement module, which divides the image data into several parallel branches. Each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, and combines a learnable offset function to achieve adaptive offset of sampling points, capturing the local detail information of the branch; uses a linear layer to capture the global information of each branch; cascades the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map;
[0107] Detector, which detects pedestrian features in the multi-receptive field feature map.
[0108] Figure 6 The internal structure diagram of a computer device in an embodiment is shown. The computer device may specifically be the Figure 1 computer device in. As Figure 6 shown, the computer device includes a processor, a memory, a network interface, and an input device connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the motion posture detection method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the motion posture detection method. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0109] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0110] In the sixth embodiment, a computer device is proposed. The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0111] Obtain image data;
[0112] Divide the image data into a number of parallel branches. Each of the branches generates a rectangular sampling grid using a different rectangular dilation rate, and combines a learnable offset function to achieve adaptive offset of the sampling points, capturing the local detail information of the branches. Use a linear layer to capture the global information of each branch. Concatenate the output feature maps of all the branches in the channel dimension to obtain a multi-receptive field feature map;
[0113] Detect pedestrian features in the multi-receptive field feature map. In the seventh embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the processor is caused to perform the following steps:
[0114] Obtain image data;
[0115] Divide the image data into a number of parallel branches. Each of the branches generates a rectangular sampling grid using a different rectangular dilation rate, and combines a learnable offset function to achieve adaptive offset of the sampling points, capturing the local detail information of the branches. Use a linear layer to capture the global information of each branch. Concatenate the output feature maps of all the branches in the channel dimension to obtain a multi-receptive field feature map;
[0116] Detect pedestrian features in the multi-receptive field feature map.
[0117] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least some of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0118] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0119] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0120] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for detecting a motion posture, characterized in that, The described motion posture detection method includes: Obtain image data; Divide the image data into several parallel branches. Each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, combines a learnable offset function to achieve adaptive offset of sampling points, and captures local detail information of the branch; uses a linear layer to capture the global information of each branch; concatenates the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map; Detect pedestrian features in the multi-receptive field feature map; Insert different numbers of hole elements between sampling points in the convolution kernel of the rectangular sampling grid to form a rectangular sampling range for guiding the network to capture pedestrian features in a rectangular form. The pedestrian features are specifically: ; Among them, and respectively represent the vertical and horizontal expansion ratios of the i-th said branch; Use a learnable offset function to achieve adaptive offset of sampling points to obtain the output feature map of the i-th branch: ; Among them, represents the output feature map of the i-th said branch, represents the position of the pixel point in the entire said image data, is the predefined position of the n-th point, represents the weight of the convolution kernel, represents the learnable offset, represents the input feature map obtained by the i-th said branch through the rectangular dilation deformable convolution mechanism; The specific operation process of using a linear layer to capture the global information of each branch is as follows: ; Among them, represents the output feature map of the $i$-th branch, BN(·) represents the batch normalization operation, and ReLU(·) represents the activation function, represents the first linear layer, represents the second linear layer, represents the feature mapping; The specific operation of concatenating the output feature maps of all branches in the channel dimension is as follows: ; Among them, represents the multi-receptive field feature map; in each of the branches, represents the feature map obtained by performing rectangular dilation deformable convolution and adaptive offset processing on the i-th branch, represents the feature map obtained by performing adaptive feed-forward neural network processing, represents the output feature map of each branch, represents the input feature map, represents that the convolutional kernel is for the two-dimensional convolutional layer, represents the concatenation operation on the feature map in the channel dimension.
2. The motion posture detection method according to claim 1, wherein The specific method for obtaining the image data is: Obtain several feature maps with different resolutions, perform unified sampling using the bilinear interpolation algorithm to obtain a common resolution, and connect them in the channel dimension to obtain the image data.
3. The motion posture detection method according to claim 1, wherein For detecting pedestrian features in the multi-receptive field feature map, a center point-based detector is used. The steps are as follows: Use the multi-receptive field feature map after dimensionality reduction as the processing target; Predict the target center point, target scale, and target center point offset respectively. Generate a target candidate box for the input image according to the predicted target center and target scale, and fine-tune the position of the target center point through the target center point offset; Construct a ground truth label for each prediction branch in the detector, and apply a two-dimensional Gaussian mask G(·) at the position of each positive sample, which is expressed as follows: ; Among them, (i, j) represents the position of the center point of the network output, K represents the number of targets in the input image, are the coordinates, width, and height of the center point of the k-th target, and the variance of the two-dimensional Gaussian distribution and are respectively proportional to the width and height of the target; Define the target scale as the width and height of the target, and the position of the k-th positive sample is assigned to the , and is also assigned to all negative sample points within a radius of 2 of the positive sample point, and other positions are all assigned 0; The central point position of the input image is mapped to the position of the output image , and the target central point offset from the true label is defined as: ; Among them, represents the floor operation, and r is the downsampling factor.
4. A motion posture detection system, characterized in that, The motion posture detection system executes the motion posture detection method according to any one of claims 1-3. The motion posture detection system includes: A backbone network to obtain image data; An adaptive feature enhancement module that divides the image data into several parallel branches. Each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, combines a learnable offset function to achieve adaptive offset of sampling points, and captures local detail information of the branch; uses a linear layer to capture the global information of each branch; concatenates the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map; A detector to detect pedestrian features in the multi-receptive field feature map.
5. A computer device, characterized in that, It includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the motion posture detection method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the processor executes the steps of the motion posture detection method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Pedestrian detection network structure based on mixed feature pyramid and mixed expansion convolution
CN110929685A
Attitude estimation method based on adaptive receptive field network and joint loss weight
CN112241726A