Target detection method based on complementary perception feature fusion mechanism

By constructing CFFDNet, using DFSC and GWF modules to extract and empower complementary features of optical and SAR images, the problems of insufficient feature extraction and poor fusion in the multimodal detection method in the prior art are solved, and the precise detection of aircraft targets and environmental adaptability are achieved.

CN120451508AActive Publication Date: 2025-08-08HARBIN INST OF TECH

Patent Information

Application Number
CN202510599149.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-10
Publication Date
2025-08-08
Estimated Expiration
2045-05-10

AI Technical Summary

Technical Problem

The prior art is difficult to fully extract the multi-level semantic features and spatial context information of optical and SAR remote sensing images, and lacks fine measurements and efficient fusion of contributions to different modal features, resulting in limited aircraft target detection performance in complex environments.

Method used

A complementary perceptual feature fusion detection network (CFFDNet) is constructed, and the complementary features of optical and SAR images are extracted and empowered by the differential feature space perception module (DFSC) and the gate generation weighted fusion module (GWF) respectively, to enhance context information and efficiently fusion, and to improve detection accuracy using the multi-scale feature pyramid structure.

Benefits of technology

It realizes accurate detection of aircraft targets in complex scenarios, improves detection accuracy and robustness, and enhances the ability to adapt to environmental interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451508A_ABST
    Figure CN120451508A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method based on a complementary perception feature fusion mechanism, and the method proposes a complementary perception feature fusion detection network, extracts the multi-level semantic features of optical and SAR modal remote sensing images in the spatial dimension, and quantitatively analyzes the complementary difference of the two modal features in aircraft target representation. On the basis, the complementary features of the two modes are fused in a self-adaptive empowerment mode, and accurate detection of the aircraft in a complex scene is achieved; a differential feature space perception complementation module is provided, deep feature differences of optical and SAR modals in spatial dimensions can be sensitively captured, and rich context information related to an aircraft in two modal feature maps is enhanced; and a gate generation weight fusion module is provided, a weight matrix consistent with the input feature dimension is calculated to measure the contribution of the optical and SAR input features to the detection task, and weighting is carried out again to effectively fuse the target information in the two modes, so that a final accurate detection result is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection and recognition, and relates to a remote sensing image target detection method, and specifically to a target detection method based on a complementary perception feature fusion mechanism. Background Art

[0002] In recent years, aircraft target detection technology based on remote sensing images has received increasing attention and is playing an increasingly important role in civilian applications. For example, aircraft target detection can provide timely and accurate information support for applications such as airport traffic management and emergency rescue response. With the rapid development of computer vision technology, deep neural network-based target detection methods (such as Faster R-CNN, YOLO, and vision transformer) have been widely used in aircraft target detection tasks due to their ability to adaptively extract high-level semantic features from input images. Visible light images and synthetic aperture radar (SAR) images are the two main data sources for aircraft target detection. Optical remote sensing images offer rich texture and color information, intuitive visual effects, and low acquisition costs, but are susceptible to environmental factors such as illumination variations and shadow interference. SAR imaging is usable around the clock and in all weather conditions, unaffected by natural environmental factors such as clouds, rain, snow, and smoke. However, its images suffer from significant noise interference and a lack of color and texture details.

[0003] In recent years, multimodal data fusion has become a research hotspot in the field of aircraft target detection. This method can overcome the limitations of single-modal data and significantly improve the environmental adaptability of aircraft detection networks. Considering the high complementarity between optical and SAR images in terms of information dimension, fusing optical and SAR image features for joint detection has become a feasible way to improve detection accuracy and algorithm robustness. In fact, researchers have proposed a variety of aircraft detection methods based on multimodal data fusion applications and have achieved certain results in improving detection accuracy and overcoming environmental interference. Among them, feature-level fusion is a widely used and promising processing strategy for constructing multimodal target detection models. It can extract information in feature dimensions from complementary modalities and effectively fuse them to obtain excellent target detection results.

[0004] Despite this, current multimodal detection methods based on feature-level fusion still face the following problems that need to be solved urgently: (1) In the feature extraction stage, existing methods find it difficult to comprehensively consider and effectively utilize the multi-level features such as structure, texture, semantics and their spatial context information of complementary modalities, resulting in insufficient extraction of complementary features; (2) When performing multimodal feature-level fusion, existing methods lack the precise measurement and efficient fusion of the contribution of different modal features, resulting in limited detection performance of the model under complex environmental interference. Summary of the Invention

[0005] To address the issues of insufficient multimodal complementary feature extraction and a lack of precise cross-modal feature contribution measurement in multimodal aircraft target fusion detection, this paper provides a target detection method based on a complementary perceptual feature fusion mechanism. This method proposes a complementary perceptual feature fusion detection network that extracts multi-layered semantic features from optical and SAR remote sensing images in the spatial dimension, quantitatively analyzes the complementary differences between the two modal features in aircraft target representation, and then adaptively weights and fuses the complementary features of the two modalities to achieve accurate aircraft detection in complex scenarios.

[0006] The purpose of the present invention is achieved through the following technical solutions:

[0007] A target detection method based on a complementary perceptual feature fusion mechanism includes the following steps:

[0008] Step 1: Construct a Complementarity-aware Feature Fusion Detection Network (CFFDNet), which takes optical and SAR images of the same scene as input and strictly aligns the optical and SAR images in the spatial dimension. Utilize the Differential Feature Spatial-aware Complementary (DFSC) module to extract optical and SAR image features and optical-SAR differential features. Simultaneously, construct a Cascade Large-kernel Spatial Perception (CLSP) structure to process the differential features and generate the corresponding differential spatial attention map. Combined with the classic residual structure, the two modal feature maps can be used to enhance the perception of the aircraft target structure while extracting rich contextual information.

[0009] Step 2: Using the extracted optical and SAR features as input, a gate-generated weighted fusion (GWF) module is used to measure the impact of the two modal features on the final aircraft detection results, reweighting and fusing the effective information.

[0010] Step 3: The weighted fused features are passed to the neck of CFFDNet, and the multi-scale feature pyramid structure is used to further integrate and enhance the feature information at different levels. The fused multi-scale features are then input into the detection head of CFFDNet to achieve fast and accurate detection of aircraft targets in optical and SAR images.

[0011] Compared with the prior art, the present invention has the following advantages:

[0012] 1. A differential feature spatial perception complementary module is proposed, which can keenly capture the deep feature differences between optical and SAR modalities in the spatial dimension and enhance the rich contextual information related to the aircraft in the feature maps of the two modalities.

[0013] 2. A gate-generated weighted fusion module is proposed to calculate a weight matrix consistent with the input feature dimension to measure the contribution of optical and SAR input features to the detection task, and re-weight them accordingly to effectively fuse the target information under the two modalities, thereby generating the final accurate detection results. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 This is a schematic diagram of the CFFDNet architecture;

[0015] Figure 2 This is the specific structure diagram of the main modules of CFFDNet;

[0016] Figure 3 A comparison chart of the detection results of single-modal and multi-modal feature fusion. DETAILED DESCRIPTION

[0017] The technical solution of the present invention is further described below with reference to the accompanying drawings, but is not limited thereto. Any modification or equivalent replacement of the technical solution of the present invention that does not depart from the spirit and scope of the technical solution of the present invention should be included in the scope of protection of the present invention.

[0018] The present invention provides a target detection method based on a complementary perception feature fusion mechanism. The method designs a CFFDNet to capture the deep feature differences between optical and SAR modal remote sensing images in the spatial dimension. First, the DFSC module is used to extract the differential features of the optical and SAR images of the same scene, enhancing the learning of aircraft target features in the feature maps of the two modalities. Then, the GWF module is used to re-weight and fuse the effective information in the two modalities, achieving efficient utilization of the complementary features between the optical and SAR modalities and improving the detection accuracy of aircraft targets. The overall architecture is as follows Figure 1 As shown, the specific steps include:

[0019] Step 1: Build CFFDNet, taking optical and SAR images of the same scene as input to achieve strict alignment of optical and SAR images in the spatial dimension; use the DFSC module (mainly composed of CLSP units and classic residual structure units) to extract optical and SAR image features and optical-SAR differential features, and simultaneously construct a CLSP structure to process the differential features to generate the corresponding differential spatial attention map. Combined with the classic residual structure, it extracts rich contextual information while enhancing the perception of the aircraft target structure in the two modal feature maps. The specific steps are as follows:

[0020] Step 1: Construct CFFDNet by using the dual-branch backbone network (Backbone), neck (Neck) and detection head (Head) based on ResNet-50, and introduce the DFSC module and GWF module to extract the optical image X OPT and SAR image X SAR The differential features of the image are weighted and fused to achieve the integration and utilization of the complementary features of optical images and SAR images. The main modules of CFFDNet are shown in the specific structure diagram. Figure 2 shown.

[0021] Step 1 and 2: Use the dual-branch backbone network of CFFDNet to analyze the optical image X of the same scene. OPT and SAR image X SAR Perform multi-level deep feature extraction and input these features into the DFSC module at the same time. Calculate the differential feature X for the input optical features and SAR features at the same level. D,OPT and X D,SAR , thereby suppressing X OPT and X SAR The shared information in the amplifies the differential information between the two, X D,OPT and X D,SAR The calculation method is shown in the following formula:

[0022] X D,OPT =X OPT -X SAR

[0023] X D,SAR =X SAR -X OPT

[0024] Step 13: Use two cascaded deep convolutions to process the obtained differential features to obtain rich aircraft target semantics and context information representations within different receptive fields X D,i , the calculation process is shown as follows:

[0025] X D,i =W i dw (X D),i=1,2

[0026] Where, X D Represents X D,OPT or X D,SAR , W i dw (·) represents the cascaded depthwise convolution operation, W1 dw (·) is the first layer of depth convolution, W2 dw (·) is the second layer of depth convolution, X D,1 represents the output of the first layer of depthwise convolution, X D,2 Represents the output of two layers of depth convolution.

[0027] Then use two 1×1 convolution kernels to D,i Convolution calculates and generates differential spatial attention maps V with different receptive fields D,1 and V D,2 , V D,i The calculation method is shown in the following formula:

[0028] V D,i =W 1×1 (X D,i ),i=1,2

[0029] Where W 1×1 (·) denotes a 1×1 convolution operation.

[0030] Step 14: Get the differential spatial attention map V D,1 and V D,2 As input, the two are concatenated in the channel dimension and averaged and max pooled to obtain the feature map M avg and M max , M avg and M max The calculation process is as follows:

[0031] M avg =CAP(concat(V D,1 ,V D,2 ))

[0032] M max =CMP(concat(V D,1 ,V D,2 ))

[0033] Where concat(·) represents the fully connected operation, CAP(·) and CMP(·) represent the average pooling and maximum pooling operations in the channel dimension, respectively.

[0034] For the obtained M avg and M maxThe concatenation is performed along the channel dimension, and a large kernel convolution layer is applied to realize the information exchange between the two. The final spatial mask M is generated by the sigmoid function. The process can be expressed as follows;

[0035] M=Sigmoid(W 1×1 (concat(M avg ,M max )))

[0036] Where Sigmoid(·) represents the Sigmoid activation function.

[0037] The two levels of differential spatial attention map V D,i They are weighted by multiplying them spatially with the spatial selection mask and fused through a 1×1 convolution kernel to generate the final differential spatial weighted attention map V D , V D The calculation method is as follows:

[0038] V D =W 1×1 (M1⊙V D,1 +M2⊙V D,2 )

[0039] Where ⊙ represents the weighted multiplication by space, V D Indicates V D,OPT or V D,SAR .

[0040] Step 15: Use the obtained differential spatial attention map V D,OPT and V D,SAR Respectively correspond to the original input optical image X OPT and SAR image X SAR Perform spatial multiplication and weighting to obtain differential features that highlight local features after spatial weighting and This process can be expressed as follows:

[0041]

[0042] Step 16: Obtain and As supplementary information, add to X SAR and X OPT In order to generate more information features, the classical residual structure is then added to the optical feature and SAR feature branches to improve the stability of the fusion module structure, and the feature-enhanced optical image is obtained. and SAR images The specific process can be expressed as:

[0043]

[0044] Where Res(·) represents the classical residual structure.

[0045] Step 2: Using the extracted optical and SAR features as input, the GWF module (mainly composed of convolution, sigmoid, and weighted sum operations) measures the impact of the two modal features on the final aircraft detection results and reweights and combines the effective information. The specific steps are as follows:

[0046] Step 21: Optical Features and SAR characteristics As input, splicing is performed along the channel dimension, and the generated combined features are sequentially fed into the 1×1 convolution kernel and the sigmoid function to obtain the weight matrix The calculation process is expressed by the following formula:

[0047]

[0048] Step 2: Use the obtained G and 1-G as the weight matrix of the optical mode and SAR mode to obtain the final gated fusion feature X GWF , complete the feature fusion of optical image and SAR image, the fused feature map X GWF The calculation process is shown in the following formula:

[0049]

[0050] Step 3: The weighted fused features are passed to the neck of CFFDNet, and the multi-scale feature pyramid structure is used to further integrate and enhance the feature information at different levels. The fused multi-scale features are then input into the detection head of CFFDNet. The detection head contains a classification subnetwork and a regression subnetwork, which are used to predict the category and accurate location information of the aircraft target respectively. The head adopts a lightweight design to ensure detection accuracy while improving computational efficiency, achieving fast and accurate detection of aircraft targets in optical and SAR images. The specific steps are as follows:

[0051] Step 31: The fusion feature X obtained in step 2 GWF Perform a series of convolution and pooling operations to generate feature maps of different scales {C1, C2, C3, C4, C5}, where C i The scale of the original image The calculation process is shown as follows:

[0052] C2=MaxPool(Conv 3×3 (X GWF ))

[0053] C i+1=MaxPool(Conv 3×3 (C i )),i∈{2,3,4}

[0054] Where, Conv n×n (·) represents the n×n convolution operation, and MaxPool(·) represents the maximum pooling operation.

[0055] Step 32: Using a top-down path, upsample the high-level semantic features through bilinear interpolation and add them to the main elements of the shallow high-resolution features to generate a multi-scale feature pyramid {P1, P2, P3, P4, P5}. The calculation process is as follows:

[0056] P5=Conv 1×1 (C5)

[0057] P i =Conv 3×3 (C i )+Upsample(P i+1 ),i∈{4,3,2}

[0058] Where Upsample(·) represents the upsampling operation.

[0059] Each layer P i Apply the spatial attention module to obtain the feature map P i ', enhance the saliency of the target area, the calculation process can be expressed as follows:

[0060]

[0061] in, Represents element-wise multiplication.

[0062] Step 33: For each layer P obtained i 'Apply a lightweight classification network and output the category score cls for each anchor point i , at the same time, for each layer P i 'Apply the regression network to predict the offset bbox of the target bounding box i .cls i and bbox i The calculation process is as follows:

[0063] cls i =Softmax(Conv 3×3 (ReLU(Conv 1×1 (P i '))))

[0064] bbox i =Conv 3×3(ReLU(Conv 1×1 (P i ')))

[0065] Where ReLU(·) represents the activation function and Softmax(·) represents the normalization function.

[0066] Steps 3 and 4: Aggregate the classification and regression results at different levels and apply non-maximum suppression (NMS) to remove redundant detection boxes to obtain the final detection results:

[0067]

[0068] Where NMS(·) represents the non-maximum suppression operation to remove duplicate results.

[0069] Step 4: Use a publicly available, high-quality public dataset to verify the algorithm performance. The specific steps are as follows:

[0070] The present invention uses the public MAR20 dataset to conduct experiments, and divides this dataset into training and test sets in a ratio of 7:3. All unimodal baseline methods and the proposed method are implemented using MMDetection. For the comparative multimodal detection model, experiments are conducted using open source code and laboratory replication. During the model training process, the batch size is set to 4, and the training phase adopts 24 epochs. The SGD optimizer is used to update the model parameters, and the initial learning rate, momentum, and weight decay are 0.005, 0.9, and 0.0001, respectively.

[0071] Figure 3 The detection performance differences between single-modal and multi-modal feature fusion are compared. Figure 3 (a) is the detection result of a single optical mode, Figure 3 (b) is the SAR single mode detection result, Figure 3 (c) is the fusion detection result of direct addition of optical and SAR modes, Figure 3 (d) is the feature-level fusion detection result proposed by the present invention, Figure 3 (e) is the ground truth value of the dataset image. It should be noted that the green, yellow, and red boxes in the figure represent correct detections, missed detections, and false alarms, respectively. The experimental results show that the proposed remote sensing image target detection method based on a complementary perceptual feature fusion mechanism can effectively integrate the complementary information of optical and SAR modalities and effectively improve the performance of aircraft target detection in complex remote sensing scenarios.

Claims

1. A target detection method based on complementary perception feature fusion mechanism, characterized by The method comprises the following steps: Step 1: Build a complementary perception feature fusion detection network, taking optical and SAR images of the same scene as input, and achieve strict alignment of optical and SAR images in the spatial dimension; A differential feature spatial perception complementary module is used to extract optical and SAR image features, as well as optical-SAR differential features. A cascaded large-kernel spatial perception structure is constructed to process the differential features to generate corresponding differential spatial attention maps. Combined with a classic residual structure, this method not only extracts rich contextual information but also enhances the perception of the aircraft target structure from the two modal feature maps. Step 2: Using the extracted optical and SAR features as input, a gate-generated weighted fusion module is used to measure the impact of the two modal features on the final aircraft detection results, reweighting and fusing the effective information. Step 3: The weighted fused features are passed to the neck of the complementary perception feature fusion detection network, and the multi-scale feature pyramid structure is used to further integrate and enhance the feature information at different levels. The fused multi-scale features are then input into the detection head of the complementary perception feature fusion detection network to achieve fast and accurate detection of aircraft targets in optical and SAR images.

2. The target detection method based on the complementary perception feature fusion mechanism according to claim 1 is characterized in that The specific steps of step one are as follows: Step 1: Construct a complementary perception feature fusion detection network based on the dual-branch backbone network, neck and detection head based on ResNet-50, and introduce the differential feature space perception complementary module and gate generation weighted fusion module to extract the optical image X OPT and SAR image X SAR The differential features of the images are weighted and fused to achieve the integration and utilization of the complementary features of optical images and SAR images; Step 1 and 2: Use the dual-branch backbone network of the complementary perception feature fusion detection network to detect the optical image X of the same scene OPT and SAR image X SAR Perform multi-level deep feature extraction and input these features into the differential feature space perception complementary module at the same time; calculate the differential feature X for the input optical features and SAR features at the same level D,OPT and X D,SAR ; Step 13: Use two cascaded deep convolutions to process the obtained differential features to obtain rich aircraft target semantics and context information representations within different receptive fields X D,i , and then use two 1×1 convolution kernels to D,i Convolution calculates and generates differential spatial attention maps V with different receptive fields D,1 and V D,2 ; Step 14: Get the differential spatial attention map V D,1 and V D,2 As input, the two are concatenated in the channel dimension and averaged and max pooled to obtain the feature map M avg and M max , for the obtained M avg and M max They are connected in series along the channel dimension, and a large kernel convolution layer is applied to realize the information exchange between the two layers. The final spatial mask M is generated by the sigmoid function, and the differential spatial attention map V of the two levels is converted into D,i They are weighted by multiplying them spatially with the spatial selection mask and fused through a 1×1 convolution kernel to generate the final differential spatial weighted attention map V D ; Step 15: Use the obtained differential spatial attention map V D,OPT and V D,SAR Respectively correspond to the original input optical image X OPT and SAR image X SAR Perform spatial multiplication and weighting to obtain differential features that highlight local features after spatial weighting and Step 16: Obtain and As supplementary information, add to X SAR and X OPT In order to generate more information features, the residual structure is then added to the optical feature and SAR feature branches to improve the stability of the fusion module structure, and the feature-enhanced optical image is obtained. and SAR images 3. The target detection method based on complementary perceptual feature fusion mechanism according to claim 1 is characterized in that The X D,OPT and X D,SAR The calculation method is shown in the following formula: X D,OPT =X OPT -X SAR X D,SAR =X SAR -X OPT X D,i The calculation process is shown as follows: X D,i =W i dw (X D ),i=1,2 Where, X D Represents X D,OPT or X D,SAR , W i dw (·) represents the cascaded depthwise convolution operation, W1 dw (·) is the first layer of depth convolution, W2 dw (·) is the second layer of depth convolution, X D,1 represents the output of the first layer of depthwise convolution, X D,2 Represents the output of two layers of depth convolution calculation; V D,i The calculation method is shown in the following formula: V D,i =W 1×1 (X D,i ),i=1,2 Where W 1×1 (·) represents a 1×1 convolution operation; M avg and M max The calculation process is as follows: M avg =CAP(concat(V D,1 ,V D,2 )) M max =CMP(concat(V D,1 ,V D,2 )) Where concat(·) represents the fully connected operation, CAP(·) and CMP(·) represent the average pooling and maximum pooling operations in the channel dimension respectively; M is represented by the following formula: M=Sigmoid(W 1×1 (concat(M avg ,M max ))) Where, Sigmoid(·) represents the Sigmoid activation function; V D The calculation method is as follows: V D =W 1×1 (M1⊙V D,1 +M2⊙V D,2 ) Where ⊙ represents the weighted multiplication by space, V D Indicates V D,OPT or V D,SAR ; and Expressed as: Where Res(·) represents the residual structure.

4. The target detection method based on complementary perceptual feature fusion mechanism according to claim 2 is characterized in that The specific steps of step 2 are as follows: Step 21: Optical Features and SAR characteristics As input, splicing is performed along the channel dimension, and the generated combined features are sequentially fed into the 1×1 convolution kernel and the sigmoid function to obtain the weight matrix Step 2: Use the obtained G and 1-G as the weight matrix of the optical mode and SAR mode to obtain the final gated fusion feature X GWF , complete the feature fusion of optical image and SAR image, and obtain the fused feature map X GWF .

5. The target detection method based on complementary perceptual feature fusion mechanism according to claim 4 is characterized in that The calculation process of G is expressed by the following formula: X GWF The calculation process is shown in the following formula:

6. The target detection method based on complementary perceptual feature fusion mechanism according to claim 4 is characterized in that The specific steps of step three are as follows: Step 31: The fusion feature X obtained in step 2 GWF Perform a series of convolution and pooling operations to generate feature maps of different scales {C1, C2, C3, C4, C5}; Step 32: Using a top-down path, upsample the high-level semantic features through bilinear interpolation and add them to the main elements of the shallow high-resolution features to generate a multi-scale feature pyramid {P1, P2, P3, P4, P5}. i Apply the spatial attention module to obtain the feature map P i ', enhance the saliency of the target area; Step 3: For each layer of feature map P obtained i 'Apply a lightweight classification network and output the category score cls for each anchor point i , at the same time, for each layer P i 'Apply the regression network to predict the offset bbox of the target bounding box i ; Steps 3 and 4: Aggregate the classification and regression results at different levels, and apply non-maximum suppression to remove redundant detection boxes to obtain the final detection result S.

7. The target detection method based on complementary perceptual feature fusion mechanism according to claim 6 is characterized in that The calculation process of C2, C3, C4, and C5 is shown in the following formula: C2=MaxPool(Conv 3×3 (X GWF )) C i+1 =MaxPool(Conv 3×3 (C i )),i∈{2,3,4} Where, Conv n×n (·) represents n×n convolution operation, MaxPool(·) represents maximum pooling operation; The calculation process of P2, P3, P4, and P5 is as follows: P5=Conv 1×1 (C5) P i =Conv 3×3 (C i )+Upsample(P i+1 ),i∈{4,3,2} Where Upsample(·) represents the upsampling operation; P i The calculation process of ' is expressed as follows: in, Represents element-wise multiplication; cls i and bbox i The calculation process is as follows: class i =Softmax(Conv 3×3 (ReLU(Conv 1×1 (P)))) bbox i =Conv 3×3 (ReLU(Conv 1×1 (P i ))) In the formula, ReLU(·) represents the activation function, and Softmax(·) represents the normalization function; The calculation process of S is as follows: Where NMS(·) represents the non-maximum suppression operation to remove duplicate results.

Citation Information

Patent Citations

  • Remote sensing image target detection method and system based on multi-modal difference complementation fusion

    CN118628930A

  • Small target detection method for driving scene

    CN119863775A

  • Single-chip sensor multi-function imaging

    US20130300836A1

  • Contextual visual-based SAR target detection method and apparatus, and storage medium

    US20230184927A1

Cited By

  • Method and device for monitoring and processing abnormal behaviors in AGV charging room

    CN121256648A