Multi-view pedestrian detection method based on multi-stage deformable disentanglement

WO2026199105A1PCT designated stage Publication Date: 2026-10-01SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/084364
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2026-10-01

Smart Images

  • Figure CN2025084364_01102026_PF_FP_ABST
    Figure CN2025084364_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed in the present invention is a multi-view pedestrian detection method based on multi-stage deformable disentanglement. The method comprises: for a target scene, acquiring a plurality of view images, and inputting same into a pedestrian detection model, wherein the pedestrian detection model performs the following: extracting pedestrian features from the plurality of view images; disentangling task-related discriminative information from the pedestrian features to obtain effective features after first disentanglement; projecting the effective features after first disentanglement to a unified ground plane coordinate system, to generate bird's eye view features; computing dynamic fusion weights based on multi-view feature semantics themselves, and then obtaining fused ground plane world features on the basis of the dynamic fusion weights; disentangling scene‑related features from the fused ground plane world features to obtain effective features after second disentanglement; and generating a ground plane pedestrian position heatmap on the basis of the effective features after second disentanglement to obtain a multi-view pedestrian detection result. The present invention significantly improves the accuracy of pedestrian detection in complex multi-view scenes.
Need to check novelty before this filing date? Find Prior Art

Description

A Multi-View Pedestrian Detection Method Based on Multi-Stage Deformable Decoupling Technical Field

[0001] This invention relates to the field of computer vision technology, and more specifically, to a multi-view pedestrian detection method based on multi-stage deformable decoupling. Background Technology

[0002] Pedestrian detection is widely used in practical applications such as autonomous driving, video surveillance, and robot navigation. In recent years, deep learning-based multi-view pedestrian detection methods have significantly alleviated the problem of missed detection caused by occlusion in monocular detection by fusing overlapping viewpoint information from multiple synchronously calibrated cameras. Specifically, in multi-view scenes with camera images from multiple perspectives, feature extraction is required. Then, homography transformation is used to project the features from each perspective onto a unified ground plane to construct a scene-level bird's-eye view (BEV) for feature fusion, and the detection accuracy is calculated based on the final fusion result. During image feature extraction, complex multi-view scenes contain a large amount of noise information, such as background texture and lighting effects. Moreover, the use of manually designed heuristic rules or single-scale weight maps in the fusion stage makes it difficult to adapt to dynamic viewpoint configurations.

[0003] To address the issue of redundant noise in the feature space, existing solutions are based on variational autoencoders (VAEs), such as β-VAEs, which enhance the decoupling capability of the latent space by introducing a hyperparameter β. However, these solutions are limited by the trade-off between reconstruction accuracy and decoupling strength. Other solutions further optimize this by employing regularization strategies, such as using a total correlation penalty term to reduce redundant information among latent variables. Still others combine this with generative adversarial networks (GANs) to achieve decoupling by maximizing the mutual information between the latent encoding and generated data.

[0004] To address the issue of varying importance relationships between different views during multi-view fusion, existing research has employed various methods, including calculating fusion weights based on camera distance, introducing image confidence to determine the importance relationships between views, or using single-view supervision results to participate in the weighted fusion of multiple views.

[0005] Analysis reveals that current technologies for single-view recognition tasks focus solely on the channel dimension of data features, neglecting the impact of spatial deformation (such as geometric distortion caused by viewpoint differences) on feature discriminativity. This impact is particularly pronounced in multi-view crowd detection scenarios. For instance, the apparent features of a target can undergo non-rigid deformation due to factors such as viewpoint projection transformations and changes in occlusion patterns. Furthermore, current multi-view fusion methods typically rely on fixed priors (such as camera distance or projection confidence) to assign weights, ignoring the semantic relationships and channel differences between multi-view features. This results in fusion weights lacking scene adaptability and struggling to adapt to dynamic viewpoint configurations.

[0006] In summary, existing technologies still need to be improved to address the problem of decreased detection accuracy in complex multi-view scenarios due to the presence of a large amount of noise information unrelated to detection in the feature space; and the problem of heavy reliance on fixed camera layout assumptions and lack of scene adaptability caused by the use of manually designed heuristic rules or single-scale weight maps in the fusion process. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide a multi-view pedestrian detection method based on multi-stage deformable decoupling. This method includes:

[0008] For the target scene, acquire multiple view images;

[0009] The multiple view images are input into a trained pedestrian detection model to obtain multi-view pedestrian detection results;

[0010] The pedestrian detection model includes a feature extraction layer, a first deformable decoupling module, a bird's-eye view feature extraction module, a scene joint attention module, a second deformable decoupling module, and a decoder. The feature extraction layer extracts pedestrian features from the multiple view images. The first deformable decoupling module decouples task-related discriminative information from the pedestrian features to obtain effective features after the first decoupling. The bird's-eye view feature extraction module projects the effective features after the first decoupling onto a unified ground plane coordinate system through homography transformation to generate multi-view aligned bird's-eye view features. The scene joint attention module calculates dynamic fusion weights based on the semantics of the multi-view features and then obtains fused features based on the dynamic fusion weights. The second deformable decoupling module decouples scene-related features from the fused features to obtain effective features after the second decoupling. The decoder generates a ground plane pedestrian location heatmap based on the effective features after the second decoupling to obtain multi-view pedestrian detection results.

[0011] Compared with existing technologies, the advantages of this invention are that it provides a multi-view pedestrian detection method based on multi-stage deformable decoupling, especially a multi-view pedestrian detection method based on multi-stage deformable decoupling and scene joint attention fusion weights. This method preserves spatial deformation information while suppressing interference from irrelevant noise information. Furthermore, during feature fusion, the view weights do not depend on the assumption of a fixed camera layout, which can better decouple features related to and irrelevant to the detection task in a multi-view environment. At the same time, it adaptively models the importance relationship of different perspectives in multi-view fusion, significantly improving the pedestrian detection accuracy in complex multi-view scenes.

[0012] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description

[0013] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.

[0014] Figure 1 is a flowchart of a multi-view pedestrian detection method based on multi-stage deformable decoupling according to an embodiment of the present invention. Detailed Implementation

[0015] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0016] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.

[0017] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.

[0018] In all the examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0019] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0020] In summary, the multi-view pedestrian detection method based on multi-stage deformable decoupling provided by this invention includes: acquiring multiple view images for a target scene; inputting the multiple view images into a trained pedestrian detection model to obtain multi-view pedestrian detection results. The pedestrian detection model performs the following steps: extracting pedestrian features from the multiple view images; decoupling task-related discriminative information from the pedestrian features to obtain effective features after the first decoupling; projecting the effective features after the first decoupling onto a unified ground plane coordinate system through homography transformation to generate multi-view aligned bird's-eye view features; calculating dynamic fusion weights based on the semantics of the multi-view features themselves, and then obtaining fused ground plane world features based on the dynamic fusion weights; decoupling scene-related features from the fused ground plane world features to obtain effective features after the second decoupling; and generating a ground plane pedestrian location heatmap based on the effective features after the second decoupling, thereby obtaining multi-view pedestrian detection results.

[0021] Specifically, the provided multi-view pedestrian detection method based on multi-stage deformable decoupling includes the following steps:

[0022] Step S110: Construct a pedestrian detection model for multi-view pedestrian detection.

[0023] A pedestrian detection model (or multi-view pedestrian detection model) can be constructed using multiple functional modules or various types of network structures. For example, the pedestrian detection model includes a feature extraction layer, a first deformable decoupling module, a bird's-eye view feature extraction module, a scene joint attention module, a second deformable decoupling module, and a decoder. The feature extraction layer extracts pedestrian features from the multiple view images; the first deformable decoupling module decouples task-related discriminative information from the pedestrian features to obtain effective features after the first decoupling; the bird's-eye view feature extraction module projects the effective features after the first decoupling onto a unified ground plane coordinate system using homography transformation to generate multi-view aligned bird's-eye view features; the scene joint attention module calculates dynamic fusion weights based on the semantics of the multi-view features themselves, and then obtains fused ground plane world features based on these dynamic fusion weights; the second deformable decoupling module decouples scene-related features from the fused ground plane world features to obtain effective features after the second decoupling; the decoder generates a ground plane pedestrian location heatmap based on the effective features after the second decoupling to obtain the multi-view pedestrian detection result.

[0024] In one embodiment, the feature extraction layer uses ResNet-18 as the backbone network to extract pedestrian features from multiple views within the target scene. A deformable decoupling module (DDM) is embedded after the feature extraction layer to decouple discriminative information related to the multi-view detection task (such as pedestrian footholds and spatial locations). The decoupled features are projected onto a unified ground plane coordinate system using homography transformation to generate multi-view aligned bird's-eye view (BEV) features. Then, a scene joint attention module (SJA) is used to calculate dynamic fusion weights, obtaining dynamic weights based on the semantics of the multi-view features themselves; these dynamic weights are scene-adaptive. Next, the fusion weights are used to calibrate the feature importance of the multi-view features, achieving feature fusion without prior constraints. Then, a second-stage deformable decoupling module (DDM) is constructed from the fused ground plane world features to further separate scene-level task-related and irrelevant features, obtaining effective features after the second decoupling. For the effective features after the second decoupling, a decoder network is constructed using deformable convolution to adaptively adjust the receptive field, generating a ground plane pedestrian location heatmap to achieve pedestrian detection.

[0025] It should be noted that the designed pedestrian detection model includes two important modules: a deformable decoupling module and a joint scene attention module. The deformable decoupling module adaptively adjusts the receptive field based on the geometric structure of the input features, thereby accurately modeling local deformation patterns in the spatial dimension and learning high-level decoupling representations relevant to the detection task from the residual features in the data latent space. The joint scene attention module generates a reasonable and effective fusion weight map without introducing any prior knowledge. This weight map can be dynamically adjusted based entirely on the semantic features of the image itself, changing with the characteristics of the actual application scenario, thus enhancing the universality of the pedestrian detection model.

[0026] The following sections will focus on embodiments of the deformable decoupling module and the joint scene attention module.

[0027] 1) Deformable decoupling module

[0028] It should be noted that the functions and execution processes of the first deformable decoupling module and the second deformable decoupling module are basically the same, and they will be referred to as deformable decoupling modules in the following text.

[0029] In one embodiment, the deformable decoupling module includes two key steps: semantic alignment of shallow features through an instance normalization (IN) layer to effectively reduce feature distribution differences between different camera perspectives; and the introduction of a deformable disentanglement mechanism to achieve channel-space collaborative feature calibration through a dual-path attention mechanism.

[0030] First, the input features Instance normalization (IN) is performed to reduce the differences in feature distributions between images captured from different camera perspectives:

[0031] Where μ(·) and σ(·) represent the mean and standard deviation calculated independently for each channel and sample / instance in the spatial dimension, that is, the mean and standard deviation calculated in the spatial dimension for each independent channel for each instance feature, β is the translation parameter, and γ is the scaling parameter.

[0032] Subsequently, the residual feature R is extracted, and the formula is defined as:

[0033] The residual feature R retains the discriminative information related to the detection task from the original features.

[0034] Furthermore, deformable convolution is used to generate spatial attention weights, which are then combined with channel attention to enhance the multi-view detection task-related features in the residual information, thereby achieving feature decoupling.

[0035] First, deformable convolution is used to capture spatial deformation-sensitive features, generating a spatial attention map, represented as: Offset = Conv offset (R) (3) Rspatial=DeformConv(R,Offset) (4)

[0036] Among them, Conv offset For the offset prediction network, output offset k is the deformable convolution kernel size, for example, set to k=3. DeformConv is the deformable convolution operation that generates a spatial attention map.

[0037] Next, channel attention vectors are generated using the Compressed Excitation (SE) module. a=σ(W2·δ(W1·GAP(R))) (5)

[0038] Where σ(·) and δ(·) are the Sigmoid function and ReLU activation function, respectively, and GAP is global average pooling. and Here are the parameters for the fully connected layer, and r is the dimensionality reduction ratio, for example, set to 16 to reduce the number of parameters.

[0039] Finally, feature decoupling is performed, and spatial attention and channel attention are fused to obtain the features R relevant to the multi-view detection task. + Features R unrelated to the multi-view detection task - , respectively represented as: R+ (:,:,c)=a c ·Rspatial⊙R(:,:,c) (6) R - (:,:,c)=(1-a c )·Rspatial⊙R(:,:,c) (7)

[0040] Among them, a c Let R(:,:,c) represent the attention weight of the c-th channel, and let ⊙ denote element-wise multiplication. This design emphasizes the complementarity of spatial and channel attention, ensuring that R... + +R - =R.

[0041] R + As a relevant feature for multi-view detection tasks, it can be viewed as a subset of representations decoupled from the original feature F. Therefore, it needs to be restored to the instance-normalized feature space. In the middle, feature reconstruction is performed, represented as:

[0042] in, It is an effective representation obtained after reconstruction by a deformable decoupled module. Compared to the original feature F, it significantly enhances the discriminative information relevant to the multi-view detection task (such as pedestrian footholds and spatial locations). Correspondingly, the task-independent residual feature R... - The reconstructed representation is invalid, which contains interference information such as background texture and lighting noise.

[0043] To guide the model in extracting high-level decoupled representations relevant to the multi-view detection task from the underlying feature space, a dual-constraint loss function can be used in subsequent model training. This function employs an adversarial optimization mechanism to drive the dynamic separation of task-related and irrelevant features. For example, this loss is derived from L... useful and L useless The composition, specifically represented as: L dual =L useful +L useless (9)

[0044] Among them, L dual L represents the double-constraint loss. useful L represents the loss of task-related features. useless This represents the loss of task-irrelevant features.

[0045] The core idea of ​​the aforementioned adversarial optimization mechanism is that the task-related decoupling representation R... + Detection performance should be significantly improved (i.e., the entropy value of predicting similar random distributions should be reduced), while the task-independent feature R - Then it is necessary to suppress detection performance (i.e., increase the prediction entropy value).

[0046] Specifically, considering that each spatial location in the multi-view detection task corresponds to a potential pedestrian location probability distribution, the prediction entropy of the feature space needs to be quantified region by region. First, the ground plane features after multi-view aggregation... Implement proportional feature scaling. Compress its spatial dimension to a smaller value using two-dimensional adaptive average pooling. The scaling factor is defined as:

[0047] Among them, s h and s w These are scaling factors, where H and W represent the height and width of the original feature, and H′ and W′ represent the height and width of the scaled feature.

[0048] This operation preserves the semantic integrity of the region through local feature aggregation while significantly reducing computational complexity, providing an efficient and robust feature subspace for subsequent entropy calculation. (The scaled features are then processed.) Local features of each spatial location (i,j) Calculate its class probability distribution using the Softmax function. And further calculate the normalized entropy value:

[0049] in, It is the normalized entropy value. Let represent the class probability distribution of the c-th channel, where C is the number of channels.

[0050] Finally, the average entropy is calculated as the constraint loss:

[0051] Among them, L useful L represents the average entropy of task-related features. useless The average entropy represents the task-independent features. Indicates scaled features Local features at each spatial location (i,j) in the dataset. Indicates task-independent features after scaling Local features at each spatial location (i,j) in the dataset. Represents task-related features after scaling Local features at each spatial location (i,j). Softplus(·) = ln(1 + exp(·)) is used as a smooth activation function to ensure the loss value is non-negative. useful By minimizing task-related features The predicted entropy enhances its discriminative power; while L useless By maximizing task-independent features The predicted entropy is used to suppress noise interference. This adversarial optimization mechanism forces the model to explicitly separate the two types of features, namely R. + Focusing on key semantic information such as pedestrian footholds and spatial location, R... - This encodes information about lighting differences and background redundancy.

[0052] 2) Joint Scene Attention Module

[0053] In one embodiment, the scene joint attention module includes single-view channel importance modeling, cross-view semantic association modeling, and scene-level dynamic weight fusion.

[0054] First, given a multi-view feature set Where N is the number of views, C is the number of channels, and H×W is the spatial resolution. For each view F m Single-view channel importance modeling is performed. For example, spatial dimensions are compressed using global average pooling (GAP), and channel weights are learned using two fully connected layers: s m =F ex (z,W)=σ(W2·δ(W1·z m (15)

[0055] Among them, z m s represents the global average pooling weight of the m-th view. m F represents the channel attention weight of the m-th view. sq (F) indicates that the global average pooling function F is applied to the multi-view feature set F. sq , z represents the global average pooling feature set of the multi-view. σ(·) is the Sigmoid function, δ(·) is the ReLU activation function. and These are learnable parameters. r is the compression ratio, for example, set to r = 16. The weighted features are:

[0056] Among them, F scale (s,F) represents the weighting function, where s represents the set of channel attention weights for multiple views.

[0057] Weighted features of each view The semantic saliency map is obtained by compressing it to a single channel using 1×1 convolution. Where N is the number of camera views.

[0058] Subsequently, a cross-view attention mechanism is used to model cross-view semantic associations:

[0059] Where, x mF represents the global average pooling feature of the m-th view. sq (S) indicates that global average pooling is applied to the multi-view semantic saliency map S, where S represents the semantic saliency map, and S... m (i,j) represents the feature value at position (i,j) in the m-th view, [x1; x2; ...; x...]. n ] represents the global average pooling feature set of multiple views, N represents the number of camera views, and c is the cross-view weight. For learnable association matrices, the Softmax function ensures that the weights of aggregations across views are normalized.

[0060] Based on the cross-view weight c, the original features are adaptively adjusted according to the scene to obtain the joint scene-level multi-view attention fusion weight.

[0061] Finally, scene-level dynamic weight fusion is performed using joint scene-level multi-view attention fusion weights. The fusion feature M is obtained by element-wise weighted summation, and is expressed as:

[0062] In summary, the scene joint attention module adopts hierarchical modeling, with weights driven by feature semantics, exhibiting dynamic adaptability. It does not rely on camera calibration or layout priors, while enhancing the collaborative expression of local details and global correlations.

[0063] Step S120: Train the constructed pedestrian detection model using the dataset until the set loss function criteria are met.

[0064] The dataset used to train the pedestrian detection model can be a self-built dataset or a public dataset, such as MultiViewX and WildTrack.

[0065] Wildtrack is a real-world dataset used to capture crowded pedestrian scenes in a 12×36 square meter outdoor area. The dataset consists of 400 frames recorded by 7 synchronously calibrated cameras. Each frame contains an average of approximately 20 pedestrians, and each location is observed by 3.74 cameras. The ground is discretized into a 480×1440 grid, with each grid cell representing a 2.5 square centimeter area. The original images have a resolution of 1080×1920, and the dataset uses a standard segmentation method with 360 training frames and 40 test frames.

[0066] MultiviewX is a synthetic dataset generated using the Unity engine, featuring high pedestrian density for controlled occlusion analysis. It covers a 16×25 square meter area captured by six virtual cameras. Each frame contains approximately 40 pedestrians, with an average of 4.41 cameras observing each location. The ground plane is quantized into a 640×1000 grid (each cell is 2.5 square centimeters). Similar to Wildtrack, it also contains 400 frames (360 for training and 40 for testing) at a resolution of 1080×1920.

[0067] During the pedestrian detection model training phase, single-view detection loss and ground plane detection loss are used for training. At the same time, the dual-constraint loss function introduced in step S110 above is added to drive the dynamic separation of task-related features and irrelevant features through an adversarial optimization mechanism.

[0068] For example, the single-view loss is used for single-view pedestrian head and foot detection and is defined as: L img =L img,det +L img,off +L img,box (twenty one)

[0069] Among them, L img,det For single-view detection loss, detection is performed by regressing the pedestrian foot occupancy map, L img,off The offset loss for a single view is used to compensate for the parts omitted during downsampling. L img,box The bounding box regression loss is used.

[0070] In one embodiment, the dual-constraint loss function L usegul By minimizing task-related features Predicting entropy is similar to calculating L. useless Then through task-independent features Calculation, defined as:

[0071] L dual 1 =L useful 1 +L useless 1 (twenty two)

[0072] L dual 2 =L useful 2 +L useless 2 (twenty three)

[0073] Among them, L dual 1L is the first-stage double-constraint loss generated during the single-view detection stage. dual 2 L represents the two-stage dual-constraint loss generated during the ground plane detection phase. useful 1 L represents the loss of task-related features in the first stage. useless 1 L represents the task-independent feature loss for the first stage. useful 2 L represents the two-stage task-related feature loss. useless 2 This represents the task-independent feature loss in the second stage.

[0074] In one embodiment, the ground plane loss is detected and decoded using the decoupled BEV fused plane features, defined as: L world =L world,det +L world,off (twenty four)

[0075] Among them, L world It is the ground plane loss, L world,det L represents the detection loss at the ground plane. world,off This represents the offset loss to the ground plane.

[0076] Finally, the overall loss function L for training the pedestrian detection model is... all Represented as: L all =L img +L dual 1 +L dual 2 +L world (25)

[0077] During the training of the pedestrian detection model, the optimized model parameters can be obtained by minimizing the overall loss function.

[0078] Step S130: For the target scene, acquire multiple view images, and then use the trained pedestrian detection model to obtain multi-view pedestrian detection results.

[0079] After training the pedestrian detection model, the trained parameter configuration can be retained for actual pedestrian detection. During model application, detection calculations are performed based on the fused features obtained from deformable decoupling to obtain the multi-view pedestrian detection results.

[0080] To further verify the effectiveness of this invention, the MultiViewX and WildTrack datasets were used for validation. Experimental results show that the combination of DDM and SJA achieves consistent and significant improvements across multiple evaluation metrics, such as multi-object detection accuracy, multi-object detection precision, accuracy, and recall. This synergistic effect highlights the complementary advantages of task-related feature decomposition and semantically driven fusion weights in complex multi-view scenarios. Furthermore, DDM and SJA can be seamlessly integrated into existing detection pipelines, providing a flexible solution for enhancing robustness in complex multi-view scenarios.

[0081] In summary, the multi-view pedestrian detection method based on multi-stage deformable decoupling and scene joint attention fusion weights provided in this invention adaptively adjusts the receptive field according to the geometric structure of pedestrian features extracted by the neural network, thereby accurately modeling local deformation patterns in the spatial dimension. It achieves feature decoupling through a dual-path attention mechanism, separating discriminative features relevant to the detection task. Simultaneously, it jointly models the importance of single-view channels and cross-view semantic correlation through a hierarchical attention mechanism, generating dynamically adaptive scene-level fusion weights. These weights are based on the semantics of the features themselves, eliminating reliance on fixed camera layouts and manual priors. Furthermore, as a single-stage framework without ROI proposals, this invention is lightweight, uses fewer parameters, and maintains competitive performance. The model employs an end-to-end training approach, eliminating the need for staged training, significantly improving training efficiency, and achieving an optimal balance between stability and accuracy.

[0082] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.

[0083] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0084] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0085] The computer program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, Python, etc., and conventional procedural programming languages ​​such as "C" or similar languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing state information from the computer-readable program instructions. This electronic circuitry can execute the computer-readable program instructions to implement various aspects of the invention.

[0086] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0087] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0088] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0089] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. It will be known to those skilled in the art that implementation in hardware, implementation in software, and implementation using a combination of software and hardware are equivalent.

[0090] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein. The scope of the invention is defined by the appended claims.

Claims

1. A multi-view pedestrian detection method based on multi-stage deformable decoupling, comprising: For the target scene, acquire multiple view images; The multiple view images are input into a trained pedestrian detection model to obtain multi-view pedestrian detection results; The pedestrian detection model includes a feature extraction layer, a first deformable decoupling module, a bird's-eye view feature extraction module, a scene joint attention module, a second deformable decoupling module, and a decoder. The feature extraction layer extracts pedestrian features from the multiple view images. The first deformable decoupling module decouples task-related discriminative information from the pedestrian features to obtain effective features after the first decoupling. The bird's-eye view feature extraction module projects the effective features after the first decoupling onto a unified ground plane coordinate system through homography transformation to generate multi-view aligned bird's-eye view features. The scene joint attention module calculates dynamic fusion weights based on the semantics of the multi-view features and then obtains fused features based on the dynamic fusion weights. The second deformable decoupling module decouples scene-related features from the fused features to obtain effective features after the second decoupling. The decoder generates a ground plane pedestrian location heatmap based on the effective features after the second decoupling to obtain multi-view pedestrian detection results.

2. The method according to claim 1, characterized in that, The first deformable decoupling module obtains the effective features after the first decoupling according to the following steps: Instance normalization is performed on the input feature F to obtain the normalized feature. Represented as: Where μ(·) and σ(·) represent the mean and standard deviation calculated in the spatial dimension for each independent channel of each instance feature, respectively, β is the translation parameter, and γ is the scaling parameter; Extracting the residual feature R, denoted as: A spatial attention map is generated based on the residual feature R, and is represented as follows: Offset=Conv offset (R) Rspatial=DeformConv(R,Offset) Among them, Conv offset It is the offset prediction network, Offset is the offset output by the offset prediction network, DeformConv represents the deformable convolution operation, and Rspatial is the generated spatial attention map. Generate the channel attention vector according to the following formula. a=σ(W2·δ(W1·GAP(R))) Where σ(·) is the Sigmoid function, δ(·) is the ReLU activation function, and GAP is global average pooling. and These are the parameters of the corresponding fully connected layer, where r is the dimensionality reduction ratio and C is the number of channels. Based on the spatial attention map Rspatial and the channel attention vectors, the task-related features R are obtained. + , is represented as: R + (:,:,c)=a c ·Rspatial⊙R(:,:,c) Among them, a c R(:,:,c) represents the attention weight of the c-th channel, R(:,:,c) represents the semantic information of the c-th channel, and ⊙ represents element-wise multiplication. The task-related features R + Restore to the feature space after instance normalization Represented as: in, It is the effective representation of the first decoupling output of the first deformable decoupling module.

3. The method according to claim 2, characterized in that, The scene joint attention module obtains the dynamic fusion weights according to the following steps: For a given set of multi-view features For view F m Perform single-view channel importance modeling to obtain channel attention weights s m , is represented as: s m =F ex (z,W)=σ(W2 δ(W1 z m )) Where N is the number of views, C is the number of channels, H×W is the spatial resolution, σ(·) is the Sigmoid function, and δ(·) is the ReLU activation function. and For learnable parameters, z m s represents the global average pooling weight of the m-th view. m F represents the channel attention weight of the m-th view. sq (F) indicates that the global average pooling function is applied to the feature set F of the multi-view, and z represents the global average pooling feature set of the multi-view. For view F m Using channel attention weights s m Perform weighted calculations to obtain view F. m Weighted features Represented as: Among them, F scale (s,F) represents the weighting function, where s represents the set of channel attention weights for multiple views; Weighted features The image is compressed to a single channel using 1×1 convolution, thus obtaining a multi-view semantic saliency map. For the semantic saliency map, a cross-view attention mechanism is used to model cross-view semantic associations, and the cross-view weight c is obtained, which is expressed as: Where, x m F represents the global average pooling feature of the m-th view. sq (S) indicates that global average pooling is applied to the multi-view semantic saliency map S, and S m (i,j) represents the feature value at position (i,j) in the m-th view, [x1; x2; ...; x...]. n [] represents the set of global average pooling features for multiple views. It is a learnable correlation matrix; Based on the cross-view weight c, calculate the dynamic fusion weight of multiple views at the joint scene level. Represented as: in, This represents the dynamic fusion weight.

4. The method according to claim 3, characterized in that, The fusion feature is obtained by element-wise weighted summation, and is expressed as: Where M represents the fusion weight.

5. The method according to claim 4, characterized in that, The overall loss function for training the pedestrian detection model is set as: L all =L img +L dual 1 +L dual 2 +L world in: L img =L img,det +L img,off +L img,box L dual 1 =L useful 1 +L useless 1 L dual 2 =L useful 2 +L useless 2 L world =L world,det +L world,off Among them, L all L represents the total loss value. img L represents the single-view loss. dual 1 L represents the double constraint loss generated during the single-view detection phase. dual 2 L represents the dual constraint loss generated during the ground plane detection stage. img,det For the detection loss of a single view, L img,off For the offset loss of a single view, L img,box For bounding box regression loss, L useful 1 L represents the task-related feature loss during the single-view detection stage. useless 1 L represents the task-independent feature loss in the single-view detection stage. useful 2 L represents the task-related feature loss during the ground plane detection stage. useless 2 L represents the task-independent feature loss during the ground plane detection stage. world L represents the ground plane loss. world,det L represents the detection loss at the ground plane. world,off This represents the offset loss to the ground plane.

6. The method according to claim 5, characterized in that, For the single-view detection stage and the ground plane detection stage, the task-related feature loss and the task-irrelevant feature loss are calculated using average entropy, and are uniformly expressed as follows: in, This is the normalized entropy value, set as follows: in, Let L represent the class probability distribution of the c-th channel. useful L represents the average entropy of task-related features. useless The average entropy of task-independent features is Softplus(·) = ln(1 + exp(·)). Yes The feature is scaled proportionally, where H′ and W′ represent the height and width of the scaled feature. It is a scaled feature Local features of spatial location (i,j) in the middle. Represents the scaled task-related features Local features of each spatial location (i,j), Indicates task-independent features after scaling Local features of each spatial location (i,j) in the dataset.

7. The method according to claim 1, characterized in that, The feature extraction layer is constructed based on a residual network.

8. The method according to claim 1, characterized in that, The datasets used to train the pedestrian detection model include MultiViewX or WildTrack.

9. A computer-readable storage medium having a computer program stored thereon, wherein, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 8.

10. A computer device comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.