Motion posture detection method and system, computer equipment and storage medium
By introducing a rectangular expansion deformable convolution mechanism and an adaptive feedforward neural network into pedestrian detection technology, multiple receptive field feature maps are generated, which solves the detection error and missed detection problems caused by pedestrian posture changes and improves the detection accuracy.
Patent Information
- Application Number
- CN202510660071.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
When existing pedestrian detection technology deals with pedestrians with different postures, large appearance changes lead to missed detection of detectors, and insufficient detection accuracy.
The rectangular expansion deformable convolution mechanism and adaptive feedforward neural network are used to capture the local details and global information of pedestrians through the rectangular sampling grid and the learnable offset function, generate multiple receptive field feature maps, and use a central point-based detector to perform pedestrian feature detection.
It improves the accuracy of pedestrian motion posture detection, reduces false detection and missed detection, and can more accurately deal with pedestrian targets of different scales and postures.
Smart Images

Figure CN120183049A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a method, a system, a computer device and a storage medium for detecting motion postures. Background Art
[0002] Pedestrian detection is an important research direction in the field of computer vision, which aims to identify and locate all pedestrians from images or videos. This technology plays an important role in many application scenarios, such as intelligent monitoring, autonomous driving vehicles, human-computer interaction interfaces, etc.
[0003] The goal of pedestrian detection is to accurately identify the presence of pedestrians from the input visual data (such as pictures or video streams captured by cameras) and determine their positions (usually in the form of bounding boxes). This involves complex algorithms and technologies to handle various challenges, including but not limited to the pose changes of pedestrians, occlusion, background clutter, and changes in lighting conditions.
[0004] With the wide application of pedestrian detection in real scenarios such as intelligent monitoring, human-computer interaction, and autonomous driving, the requirement for detection accuracy is getting higher and higher. However, in existing pedestrian detection technologies, the appearance of pedestrians with different postures varies greatly, and the problem of pedestrian pose changes often leads to misdetection and missed detection by detectors. To accurately capture pedestrian targets in images or videos, there is an urgent need for a motion posture detection method with higher detection accuracy. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a method for detecting motion postures, aiming to improve the accuracy of pedestrian motion posture detection and avoid misdetection and missed detection caused by pedestrian pose changes.
[0006] The embodiments of the present invention are implemented as follows. A method for detecting motion postures, the method for detecting motion postures includes: Obtain image data; Divide the image data into several parallel branches. Each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, combines a learnable offset function to achieve adaptive offset of sampling points, and captures local detail information of the branch; use a linear layer to capture global information of each branch; concatenate the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map; Detect pedestrian features in the multi-receptive field feature map.
[0007] Further, the specific method for obtaining the image data is: obtain several feature maps with different resolutions, uniformly sample them using the bilinear interpolation algorithm to obtain a common resolution, and connect them in the channel dimension to obtain the image data.
[0008] Further, different numbers of hole elements are inserted between the sampling points of the rectangular sampling grid in the convolution kernel to form a rectangular sampling range for guiding the network to capture pedestrian features in a rectangular shape. The pedestrian features are specifically: ; where and respectively represent the vertical and horizontal dilation rates of the i-th branch.
[0009] Further, an adaptive offset of the sampling points is achieved using a learnable offset function to obtain the output feature map of the i-th branch: ; where represents the output feature map of the i-th branch, represents the position of the pixel point in the entire image data, is the predefined position of the n-th point, represents the weight of the convolution kernel, represents the learnable offset, represents the input feature map obtained by the i-th branch through the rectangular dilation deformable convolution mechanism.
[0010] Further, the global information of each branch is captured using a linear layer. The specific operation process is as follows: ; where represents the output feature map of the i-th branch, BN(·) represents the batch normalization operation, ReLU(·) represents the activation function, represents the first linear layer, represents the second linear layer.
[0011] Further, the output feature maps of all branches are concatenated in the channel dimension. The specific operation is as follows: ; where represents the multi-receptive field feature map; in each branch, represents the feature map of the i-th branch after rectangular dilation deformable convolution and adaptive offset processing, represents the feature map after adaptive feed-forward neural network processing, represents the output feature map of each branch, represents the input feature map, represents a two-dimensional convolutional layer with a convolution kernel of and represents the concatenation operation of the feature maps in the channel dimension.
[0012] Further, to detect pedestrian features in the multi-receptive field feature map, a center-point based detector is used, and the steps are as follows: Use the multi-receptive field feature map after dimensionality reduction as the processing target; Predict the target center point, target scale, and target center point offset respectively. Generate a target candidate box for the input image according to the predicted target center and target scale, and fine-tune the position of the target center point through the target center point offset; Construct a ground truth label for each prediction branch in the detector, and apply a two-dimensional Gaussian mask G(∙) at the position of each positive sample, which is expressed as follows: ; where (i,j) represents the position of the center point output by the network, K represents the number of targets in the input image, are the coordinates, width, and height of the center point of the k-th target, and the variance and are respectively proportional to the width and height of the target; Define the target scale as the width and height of the target. The position of the k-th positive sample is assigned to the of the k-th target, and is also assigned to all negative sample points within a radius of 2 from the positive sample point, and other positions are all assigned 0; The position of the center point of the input image is mapped to the position in the output image, and the ground truth of the target center point offset is defined as: ; where represents the floor operation, that is, taking the largest integer not greater than this value, and r is the downsampling factor.
[0013] Another object of the embodiment of the present invention is a motion posture detection system, which is characterized in that the motion posture detection system executes the motion posture detection method, and the motion posture detection system includes: A backbone network to obtain image data; An adaptive feature enhancement module that divides the image data into several parallel branches. Each of the several branches uses different rectangular dilation rates to generate rectangular sampling grids, combines a learnable offset function to achieve adaptive offset of sampling points, and captures local detail information of the branches; uses a linear layer to capture global information of each branch; concatenates the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map; A detector that detects pedestrian features in the multi-receptive field feature map.
[0014] Another object of the embodiments of the present invention is a computer device, characterized by including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the processor executes the steps of the motion pose detection method.
[0015] Another object of the embodiments of the present invention is a computer-readable storage medium, characterized in that a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the processor executes the steps of the motion pose detection method.
[0016] A motion pose detection method provided by the embodiments of the present invention first proposes a rectangular dilated deformable convolution mechanism, which combines the advantages of the wide receptive field range of dilated convolution and the adaptive sampling ability of deformable convolution. To avoid the limitation of the local receptive field range of the method of the present invention and capture global features, an adaptive feed-forward neural network is designed. On the basis of the previous step, the global information is modeled through a linear layer, significantly expanding the feature representation ability. On this basis, the present invention further proposes a star mechanism, which divides the image data into several independent branches to achieve the sequential cascade of the rectangular dilated deformable convolution mechanism module and the adaptive feed-forward neural network within each branch, and at the same time makes full use of the local detail capture ability and the global information modeling ability, so that the method can accurately and parallelly process targets of different scales.
[0017] The present invention is mainly embodied in several aspects: First, different numbers of hole elements are inserted between sampling points, and an innovative rectangular sampling grid scheme is designed to guide the network to pay more attention to rectangular pedestrian targets; Second, on the basis of the rectangular sampling grid, deformable convolution is further combined, and an innovative rectangular dilated deformable convolution (RDDC) mechanism is proposed to achieve adaptive sampling within the rectangular receptive field range to cope with the complex pose changes of pedestrians; Third, the Stars mechanism is designed to construct an AFE-Module, and multi-scale feature maps are processed through a four-branch parallel structure, and finally a multi-receptive field feature map is obtained. Description of the Drawings
[0018] Figure 1 It is an application environment diagram of the motion pose detection method provided by the embodiments of the present invention; Figure 2 It is a pedestrian detection network (MAFE-Net) based on multi-receptive field adaptive feature enhancement provided by the embodiments of the present invention; Figure 3The structural diagram of the Adaptive Feature Enhancement Module (AFE-Module) provided by the embodiments of the present invention; Figure 4 The example diagram of the Rectangular Dilation Deformable Convolution Block (RDDC-Block) provided by the embodiments of the present invention; Figure 5 The schematic diagram showing the changes in the receptive field range and sampling point positions of the four sub-branches and their fused feature maps provided by the embodiments of the present invention; Figure 6 The internal structure block diagram of a computer device in one embodiment. Detailed implementation manners
[0019] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0020] Figure 1 The application environment diagram of a motion posture detection method provided by the embodiments of the present invention, as Figure 1 shown, in this application environment, it includes an image acquisition device and a computer device.
[0021] The image acquisition device refers to a hardware tool for capturing, recording or generating digital images or videos, which can be a camera, a scanner, a digital camera or an industrial camera, etc.
[0022] The computer device can be a smart phone, a tablet computer, a notebook computer or a desktop computer, or a cloud server providing basic cloud computing services such as cloud servers, cloud databases, cloud storage and CDN. The image acquisition device and the computer device can be connected through a network, and the present invention does not make any restrictions here.
[0023] As Figure 2 and Figure 3 shown, in the first embodiment, a motion posture detection method is proposed. In this embodiment, the method is mainly applied to the above-mentioned Figure 1 computer device for illustration. A motion posture detection method may specifically include the following steps: Step S202, obtaining image data.
[0024] In this embodiment, as Figure 2As shown, the required image data is obtained in the backbone network (HRNet). The image acquisition device inputs a series of images into the backbone network, and the image data can be obtained only after being processed by the backbone network. In this embodiment, the image acquisition device inputs a 3-channel image with a size of H×W, and performs feature extraction on the image through the backbone network. The backbone network is a parallel structure, consisting of four branches with different resolutions. By parallelly processing multiple resolution branches, the backbone network can combine the powerful semantic information of the feature map with accurate position information, and promote information interaction between different branches. The backbone network finally outputs four feature maps, and the pixel ratios are 4 times, 8 times, 16 times, and 32 times the input image ratio respectively. The feature map with a high downsampling rate contains rich deep semantic information, but due to its low resolution and relatively lack of details, it is not ideal for detection. In addition, considering that small-scale pedestrian targets account for a relatively large proportion in the pedestrian detection task, in this embodiment, the bilinear interpolation algorithm is used to uniformly upsample these four output feature maps to a common resolution and concatenate them in the channel dimension to fully integrate multi-scale information.
[0025] Step S204: Divide the image data into several parallel branches, and each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, and realizes the adaptive offset of the sampling points in combination with a learnable offset function to capture the local detail information of the branch; use a linear layer to capture the global information of each branch; concatenate the output feature maps of all the branches in the channel dimension to obtain a multi-receptive field feature map.
[0026] In this embodiment, as Figure 2 shown, the method adopted is a pedestrian detection network (MAFE-Net) based on multi-receptive field adaptive feature enhancement, aiming to enhance the pedestrian features in the feature map from the perspective of feature extraction and make the network pay more attention to pedestrian targets. Specifically, as Figure 3As shown, a rectangular dilated deformable convolution (RDDC) mechanism is first proposed. On this basis, a rectangular dilated deformable convolution module (RDDC-Block) is constructed, which combines the advantages of the wide receptive field range of dilated convolution and the adaptive sampling ability of deformable convolution. RDDC first uses the idea of inserting hole elements between sampling points in dilated convolution and designs a unique rectangular sampling grid scheme based on the common aspect ratio of pedestrians, thereby expanding the sampling range to be closer to the actual bounding box of pedestrians. Then, on the basis of this sampling grid scheme, an adaptive offset function of deformable convolution is introduced, enabling the sampling points to be dynamically adjusted according to the input feature map of different pedestrian postures. However, relying solely on the RDDC-Block is still limited by the local receptive field range and cannot fully capture global features. Therefore, the present invention designs an adaptive feed-forward neural network (AFFN-Block), which models global information through linear layers, significantly expanding the feature representation ability of the RDDC-Block. On this basis, the present embodiment proposes a star-shaped (Stars) mechanism and constructs an adaptive feature enhancement module (AFE-Module). The Stars mechanism uses four independent sub-branches to achieve the sequential cascade of the RDDC-Block and the AFFN-Block within each branch, while fully utilizing the local detail capture ability of the RDDC-Block and the global information modeling ability of the AFFN-Block, enabling the AFE-Module to process targets of different scales in parallel.
[0027] Step S206, detecting pedestrian features in the multi-receptive field feature map.
[0028] In this embodiment, the multi-receptive field feature map is detected in the detector. MAFE-Net uses a center-point-based detector. After the feature map containing fine-grained target information enters the detector, it first undergoes a 3×3 Conv layer for dimensionality reduction, reducing the channel dimension to 256. Then it passes through three parallel 1×1 Conv layers, predicting the target center point (center), target scale (scale), and target center point offset (offset) respectively. Finally, target candidate boxes of the input image are generated based on the predicted target center and target scale, and the center point position is fine-tuned through the target center point offset information, thereby further improving the performance of the detector.
[0029] In the embodiment, it is necessary to construct ground truth labels for each prediction branch in the detector. For the ground truth label of the target center point, since it is difficult to determine the accurate center point position of a target, in order to reduce the uncertainty of a large number of negative samples around the positive sample, a two-dimensional Gaussian mask G(∙) is applied at the position of each positive sample, and the formula is expressed as follows: ; Among them, (i, j) represents the position of the center point output by the network, K represents the number of targets in the input image, are the coordinates, width, and height of the center point of the k-th target, and the variance of the two-dimensional Gaussian distribution and are proportional to the width and height of the target respectively; Define the target scale as the width and height of the target. For its true label, the position of the k-th positive sample is assigned to the of the k-th target. To reduce the error caused by point prediction, all negative sample points within a radius of 2 from the positive sample point are also assigned, and other positions are assigned 0; The position of the center point of the input image is mapped to the position of the output image. The target center point offset true label is defined as: ; Among them, represents the floor operation, that is, taking the largest integer not greater than this value, and r is the downsampling factor.
[0030] In the second embodiment, step S202 specifically includes steps S302 to S308. As Figure 3 shown, steps S302 to S308 of this embodiment are executed by the AFE-Module, and the AFE-Module is composed of the RDDC-Block and the AFFN-Block; the RDDC-Block executes steps S302 and S304, the AFFN-Block executes step S306, and step S308 is about the Stars mechanism of this method.
[0031] Step S302, insert different numbers of hole elements between the sampling points in the convolution kernel of the rectangular sampling grid to form a rectangular sampling range for guiding the network to capture the pedestrian features of the rectangular shape. The pedestrian features are specifically: ; ; Among them, and represent the vertical and horizontal dilation rates of the i-th branch respectively. It should be noted that when and are both set to 1, they correspond to the standard grid R(1, 1) of the standard convolution. As Figure 2 shown, different dilation rates and , the dilation rates are (8, 4), (6, 3), (4, 2), and (2, 1) respectively. These dilation rates are called rectangular dilation rates (RDRs), which are selected to achieve a fixed ratio dilation of 2 to 1 in the vertical and horizontal dimensions.
[0032] In this embodiment, to reduce the number of parameters and computational cost of the entire module, a 1×1 convolutional layer is first applied to the output feature map from the backbone network , thereby obtaining the input feature map of the RDDC-Block : ; Pedestrians appear in the image, and regardless of their size or resolution, their predicted bounding boxes always appear as rectangles. Since the aspect ratio of most detection boxes is approximately 2 to 1, this method proposes a rectangular dilation deformable convolution (RDDC) mechanism to extract pedestrian features, which improves the flexibility of the network in processing pedestrian postures. The implementation of RDDC is as Figure 3 shown. Constrained by the standard convolutional receptive field range limitation and the inflexibility of sampling points, a rectangular sampling grid is first generated using the rectangular dilation rate (RDR) , realizing the expansion of the rectangular receptive field of the traditional convolutional operator. Then, a learnable offset function is introduced to achieve the adaptive offset of the sampling points. The advantage of this mechanism is that it effectively reduces the interference of background information and enables the network to focus on the features of pedestrian targets.
[0033] Step S304, use the learnable offset function to achieve the adaptive offset of the sampling points, and obtain the output feature map of the i-th branch: ; Among them, represents the output feature map of the i-th branch, represents the position of the pixel point in the entire image data, is the predefined position of the n-th point, represents the weight of the convolutional kernel, represents the learnable offset, represents the input feature map obtained by the i-th branch through the rectangular dilation deformable convolution mechanism.
[0034] In this embodiment, a learnable deformable convolution sampling point offset is introduced, which can flexibly adjust the sampling point position. For the learnable sampling point offset in the 3×3 convolution operation, each sampling point has an offset value in the vertical and horizontal directions. Since there are a total of 9 sampling points, this will generate 18 offset values in the vertical and horizontal directions. To obtain the learnable offset , a 3×3 standard convolution with 18 output channels is specifically designed, as shown in Figure 2 the RDDC-Block of . This convolution effectively obtains these offsets, thus realizing the adaptive adjustment of sampling points. Finally, the weights
[0035] are weighted and summed with the pixel values corresponding to the offset sampling points to obtain the output feature map of the i-th branch. This innovative design not only effectively solves the sparse sampling problem caused by dilated convolution, but also improves the flexibility of convolution sampling. ; wherein, represents the output feature map of the i-th branch, BN(·) represents the batch normalization operation, ReLU(·) represents the activation function, represents the first linear layer, represents the second linear layer.
[0036] Although the RDDC-Block effectively expands the receptive field, it is still limited in modeling long-range global information. To solve this problem, this embodiment designs a new AFFN-Block, which consists of a linear layer and an activation function and is specifically used to capture global information. Through this design, the network can not only capture local fine-grained information, but also establish the dependency relationship between global information.
[0037] As shown in the AFFN-Block of Figure 3 , first, BatchNorm (BN, batch normalization) and the ReLU function are performed on the input . Then, the first linear layer is applied to the feature map , and the number of output neurons is four times the number of input neurons, which helps the network learn more complex feature representations. Subsequently, the ReLU activation function is used to introduce non-linear features into the network. Finally, through the second linear layer , the number of neurons is restored to the number of input neurons of , ensuring the effective integration of information. Due to the information interaction between neurons in the linear layer, the AFFN-Block can integrate the information in the entire spatial dimension of the input features, thereby capturing global information.
[0038] Step S308, concatenate the output feature maps of all the branches in the channel dimension, and the specific operation is as follows: ; wherein, Represents the multiple receptive field feature map; in each of the branches, represents the feature map of the i-th branch after rectangular dilation deformable convolution and adaptive offset processing, represents the feature map processed by the adaptive feedforward neural network, represents each branch output feature map, represents the input feature map, The convolution kernel is The two-dimensional convolutional layer, Indicates the cascade operation of feature maps in the channel dimension.
[0039] In this embodiment, the Stars mechanism can further capture the multiple receptive fields of semantic information. Figure 3 As shown in the AFE-Module of Fig. 1, a four-branch parallel structure is first constructed, in which the RDDC-Block and AFFN-Block are sequentially embedded in each branch. Different RDRs are applied to each branch to expand the range of the receptive field and enrich the feature diversity. The specific RDRs used in each branch are: (8, 4), (6, 3), (4, 2) and (2, 1). Then all branches are connected in the channel dimension to achieve effective multi-scale information fusion and obtain multiple receptive field feature maps. Through the above multi-scale design, although the aspect ratio of some pedestrian detection boxes is not 2 to 1, the AFE-Module can still process them by fusing complementary information from different sub-branches.
[0040] Figure 4 This is an example diagram of RDDC-Block when RDR is (8, 4). Figure 4 (a) shows the sampling point location and receptive field range of the 3 × 3 standard convolution; Figure 4 (b) in the figure shows the expansion process of the receptive field; Figure 4 (c) in the figure represents the offset process of the sampling points; Figure 4 The boxed areas in (b) and (c) represent the extended rectangular receptive field range. Figure 4 (b) The matrix in represents the sampling grid . Figure 5 Schematic diagram showing the changes in the receptive field range and sampling point positions of the four sub-branches and their fused feature maps in the Stars mechanism. Figure 5 (a) to (d) represent the four sub-branches of the RDDC-Block with RDRs of (8, 4), (6, 3), (4, 2) and (2, 1), respectively; Figure 5 (e) in the figure shows the change of the multiple receptive field ranges of the fused feature map. Figure 5Shows the changes in the sampling point positions and receptive field ranges of each sub-branch. This design enhances feature representation while ensuring the integrity and accuracy of pedestrian information with different scales and poses.
[0041] Finally, the output feature maps of all branches are concatenated in the channel dimension to form a feature map with multiple receptive fields and rich deep semantic information. Then, by using a 1×1 convolutional layer adjust the number of channels and perform a residual connection with the input feature map to obtain the final output feature map . This design not only speeds up the network convergence rate but also improves the generalization of the AFE-Module.
[0042] In the third embodiment, the loss function of this method is calculated based on the first and second embodiments.
[0043] The total loss function in this application consists of three parts: the target center point loss , the target scale loss and the target center point offset loss . Since the target center points in the input image are fewer than the non-target center points, it causes an imbalance between positive and negative samples, which is not conducive to the training of the network model. Therefore, Focal Loss is used to solve the problem of imbalance between positive and negative samples. For this purpose, the center point classification loss function is defined as: ; where K represents the number of targets in an image, W and H represent the width and height of the image respectively, represents the probability that the predicted coordinate point (i,j) belongs to the target center point, β and γ are two hyperparameters, which are set to β = 4 and γ = 2 respectively in this application, represents the Gaussian heat map of the (i,j) coordinate point.
[0044] The SmoothL1 loss function is used to calculate the target scale loss and the target center point offset loss, and its formula is: ; where, represent the predicted value and the true value of the k-th target scale respectively, represent the predicted value and the true value of the k-th target center point offset respectively.
[0045] In summary, the total loss function L is defined as: ; where, They respectively represent the weights of the target center point loss, scale loss, and center point offset loss, which are respectively set to 0.01, 1, and 0.1 in this application.
[0046] In the fourth embodiment, based on the methods of the above three embodiments, specific experiments and results are given.
[0047] The MAFE-Net algorithm proposed in this embodiment is implemented using PyTorch and runs on four A100 PCIE-40GB-GPU devices. The backbone network HRNet of this algorithm uses the model weights pre-trained on the ImageNet dataset. In addition, for the object detection task mentioned in this application, the Adam optimization algorithm is adopted. On the CityPersons dataset for pedestrian detection in object detection, the size of the input image is set to 640×1280, the number of iterations in the training stage is set to 150, the batch size is set to 16, and the initial learning rate is set to , and the learning rate is multiplied by 0.1 every 50 iterations. This embodiment uses the average miss rate as an evaluation metric on the CityPersons dataset.
[0048] The evaluation subsets in CityPersons are divided into low occlusion and multi-scale subsets to more fairly compare the effectiveness of methods in dealing with multi-pose pedestrians. On the one hand, for targets with pixel values greater than 50, those with target visibility ranges in [0.65, 1] are classified under the reasonable occlusion subset (Reasonable), those with ranges in [0.65, 0.9] are classified under the partial occlusion subset (Partial), and those with ranges in [0.9, 1] are classified under the bare occlusion subset (Bare). On the other hand, for targets with visibility greater than 0.65, those with pixel value ranges in [50, 75] are classified as the small-scale subset (Small), those with ranges in [75, 100] are classified as the medium-scale subset (Medium), and those with ranges greater than 100 are classified as the large-scale subset (Large).
[0049] Table 1 presents the experimental results of this embodiment on the CityPersons dataset and compares the average miss detection rate with that of existing detection algorithms in dealing with multiple pedestrian postures. The average miss detection rate of the MAFE-Net algorithm on all low-occlusion and multi-scale subsets is lower than that of other algorithms. Especially on the reasonable occlusion subset (Reasonable) and small-scale subset (Small), the average miss detection rates are 9.0% and 10.3% respectively, achieving improvements of 1.4% and 3.9% compared with the CSP algorithm using the same backbone network (HRNet). Even when compared with the PEN method also used to handle pedestrian posture problems, MAFE-Net also shows strong competitiveness. The experimental results fully demonstrate the effectiveness of the adaptive feature enhancement module in MAFE-Net for multi-posture pedestrian detection.
[0050] Table 1 Comparison of the average miss detection rate of MAFE-Net with existing methods on CityPersons
[0051] In the fifth embodiment, a motion posture detection system is provided, characterized in that the motion posture detection system includes the steps and sub-steps of the motion posture detection method described in the first to third embodiments, and the motion posture detection system includes: A backbone network for acquiring image data; An adaptive feature enhancement module that divides the image data into several parallel branches. Each of the branches generates a rectangular sampling grid using different rectangular dilation rates, realizes the adaptive offset of the sampling points in combination with a learnable offset function, and captures the local detail information of the branches; uses a linear layer to capture the global information of each branch; cascades the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map; A detector for detecting pedestrian features in the multi-receptive field feature map.
[0052] Figure 6 The internal structure diagram of a computer device in an embodiment is shown. The computer device can specifically be Figure 1 the computer device in. As Figure 6As shown, the computer device includes a processor, a memory, a network interface, and an input device connected via a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement a motion posture detection method. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the motion posture detection method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0053] Those skilled in the art can understand that Figure 6 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0054] In the sixth embodiment, a computer device is proposed. The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: Obtain image data; Divide the image data into several parallel branches. Each of the branches generates a rectangular sampling grid using a different rectangular dilation rate, combines a learnable offset function to achieve adaptive offset of the sampling points, captures the local detail information of the branch; uses a linear layer to capture the global information of each branch; concatenates the output feature maps of all the branches in the channel dimension to obtain a multi-receptive field feature map; Detect pedestrian features in the multi-receptive field feature map. In the seventh embodiment, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the processor is caused to execute the following steps: Obtain image data; Divide the image data into several parallel branches. Each of the branches generates a rectangular sampling grid using a different rectangular dilation rate, combines a learnable offset function to achieve adaptive offset of the sampling points, captures the local detail information of the branch; uses a linear layer to capture the global information of each branch; concatenates the output feature maps of all the branches in the channel dimension to obtain a multi-receptive field feature map; Detect the pedestrian features in the multi-receptive field feature map.
[0055] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0056] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0057] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0058] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for detecting a motion posture, characterized in that, The described motion posture detection method includes: Obtain image data; Divide the image data into several parallel branches. Each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, combines a learnable offset function to achieve adaptive offset of sampling points, and captures local detail information of the branch; uses a linear layer to capture the global information of each branch; concatenates the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map; Detect pedestrian features in the multi-receptive field feature map.
2. The method for detecting a motion posture according to claim 1, characterized in that, The specific method for obtaining the image data is as follows: Obtain several feature maps with different resolutions, perform unified sampling using the bilinear interpolation algorithm to obtain a common resolution, and connect them in the channel dimension to obtain the image data.
3. The method for detecting a motion posture according to claim 1, characterized in that, The rectangular sampling grid inserts different numbers of hole elements between sampling points in the convolution kernel to form a rectangular sampling range for guiding the network to capture pedestrian features in a rectangular form. The pedestrian features are specifically: ; Among them, and respectively represent the vertical and horizontal expansion ratios of the i-th said branch.
4. The method for detecting a motion posture according to claim 3, characterized in that, Use a learnable offset function to achieve adaptive offset of sampling points to obtain the output feature map of the i-th branch: ; Among them, represents the output feature map of the i-th branch, represents the position of the pixel point in the entire image data, is the predefined position of the n-th point, represents the weight of the convolutional kernel, represents the learnable offset, represents the input feature map obtained by the i-th branch through the rectangle dilation deformable convolution mechanism.
5. The method for detecting a motion posture according to claim 4, characterized in that, The operation process of using a linear layer to capture the global information of each branch is as follows: ; Among them, represents the output feature map of the i-th branch, BN(·) represents the batch normalization operation, and ReLU(·) represents the activation function, represents the first linear layer, represents the second linear layer, represents the feature mapping.
6. The method for detecting a motion posture according to claim 5, characterized in that, The operation of concatenating the output feature maps of all branches in the channel dimension is as follows: ; Among them, represents the multi-receptive field feature map; in each of the branches, represents the feature map of the i-th branch after rectangular dilation deformable convolution and adaptive offset processing, represents the feature map processed by the adaptive feed-forward neural network, represents the output feature map of each branch, represents the input feature map, represents that the convolutional kernel is of the two-dimensional convolutional layer, represents the concatenation operation on the feature map in the channel dimension.
7. The method for detecting a motion posture according to claim 1, characterized in that, For detecting pedestrian features in the multi-receptive field feature map, a center point-based detector is used. The steps are as follows: Use the multi-receptive field feature map after dimensionality reduction as the processing target; Predict the target center point, target scale, and target center point offset respectively. Generate a target candidate box for the input image according to the predicted target center and target scale, and fine-tune the position of the target center point through the target center point offset; Construct a ground truth label for each prediction branch in the detector, and apply a two-dimensional Gaussian mask G(∙) at the position of each positive sample, expressed as follows: ; Among them, (i, j) represents the position of the center point of the network output, K represents the number of targets in the input image, are the coordinates, width, and height of the center point of the k-th target, and the variance of the two-dimensional Gaussian distribution and are respectively proportional to the width and height of the target; Define the target scale as the width and height of the target, and the position of the k-th positive sample is assigned to the k-th target's , and is also assigned to all negative sample points within a radius of 2 from the positive sample point, and other positions are all assigned 0; The central point position of the input image is mapped to the position of the output image , and the target central point offset from the true label is defined as: ; Among them, represents the floor operation, and r is the downsampling factor.
8. A system for detecting a motion posture, characterized in that, The motion posture detection system executes the motion posture detection method according to any one of claims 1-7. The motion posture detection system includes: A backbone network to obtain image data; An adaptive feature enhancement module that divides the image data into several parallel branches. Each of the several branches generates a rectangular sampling grid using different rectangular dilation rates, combines a learnable offset function to achieve adaptive offset of sampling points, and captures local detail information of the branch; uses a linear layer to capture the global information of each branch; concatenates the output feature maps of all branches in the channel dimension to obtain a multi-receptive field feature map; A detector to detect pedestrian features in the multi-receptive field feature map.
9. A computer device, characterized in that It includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the motion posture detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the processor executes the steps of the motion posture detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Pedestrian detection network structure based on mixed feature pyramid and mixed expansion convolution
CN110929685A
Attitude estimation method based on adaptive receptive field network and joint loss weight
CN112241726A
Action detection by exploiting motion in receptive fields
US20190354835A1
Target detection method and apparatus
WO2021098261A1