A Visual Object Detection Method and Device Based on a Hybrid Convolution Residual Structure

By adopting a hybrid convolutional residual structure and category balance loss function in visual detection technology, the detection performance of obstacles at different distances of 0-20m in tire hanging scenes is improved, and the problem of poor detection performance in the prior art is solved, and higher detection accuracy and real-time performance are achieved.

CN114913414BActive Publication Date: 2025-06-13NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210423216.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-06-13
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

The existing visual detection technology has poor detection performance of obstacles at different distances from 0-20m in tire crane scenarios, affecting the early warning effect.

Method used

A visual object detection method based on a hybrid convolutional residual structure is adopted, a hybrid cavity convolutional residual network HDResNet is used as the backbone network, and a category balanced loss function BLoss is designed to improve detection performance.

Benefits of technology

On the basis of ensuring the close-range detection effect, the detection capabilities of medium and long distances are significantly enhanced, the overall detection capabilities of tire cranes are improved to meet the real-time requirements of tire cranes for collision prevention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913414B_ABST
    Figure CN114913414B_ABST
Patent Text Reader

Abstract

A visual target detection method and device based on a hybrid convolutional residual structure. A visual detection model is constructed using the target detection network CenterNet, and the hybrid dilated convolutional residual network HDResNet is used as the backbone network of the visual detection model. An image training set is used to train the visual detection model, and the images in the training set include targets at different visual distances. The trained model is used to perform visual detection on targets at different visual distances in the input image simultaneously. The visual target detection method based on hybrid convolution of the present invention realizes anti-collision for rubber-tyred gantry cranes and has high comprehensive detection accuracy and detection speed at near, medium, and far distances.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine vision detection, relates to the safe operation of a rubber-tyred gantry crane, and specifically relates to a visual target detection method and device based on a hybrid convolutional residual structure. Background Art

[0002] The rubber-tyred gantry crane is a large-scale crane equipment for loading and unloading containers at ports, abbreviated as RTG. The conventional safety anti-collision automation scheme for RTGs is realized based on lidar target detection technology. The lidar automatically detects obstacles and detects pedestrians and truck obstacles on the driving track of the RTG. However, in rainy and foggy weather environments, the detection performance of lidar target detection technology decays severely, resulting in system failure. Based on vision-based target detection technology, the material cost is low and the deployment is simple. However, the existing vision detection technology has unstable detection of targets at a distance of 5-10 m in the RTG scenario and poor detection ability for targets at a distance of 10-20 m. In response to this problem, many methods have been proposed by relevant scholars to improve the target detection ability. Most of them focus on improving the detection ability of large targets at close range by optimizing the network structure. There are also some methods for detecting small targets at long range, such as FPN to establish a multi-scale feature pyramid to retain the features of small targets, and the Scale Match[1] method to improve the micro-target scenario through multi-scale pre-training. However, there is no solution that can comprehensively improve the detection ability at both long and short distances. Starting from the direction of the convolutional kernel, the present invention enhances the detection ability at medium and long distances on the basis of ensuring the detection effect at close range by means of hybrid dilated convolution to increase the receptive field of the neural network, thereby improving the overall detection ability of the RTG at 0-20 m.

[0003] Traditional convolutional neural networks increase the network depth by stacking convolutional and pooling layers to improve target detection performance. Among them, the pooling layer reduces the image size, and subsequent upsampling must be used to restore the original image size for further processing. Some pixel information is lost during the pooling and upsampling processes, thereby reducing the detection accuracy. Yu[2] et al. proposed a dilated convolution model that has the ability to increase the receptive field and improve network performance without changing the number of parameters.

[0004] The lightweight model CenterNet is selected for the anti-collision visual detection of the RTG, which has excellent detection performance for large targets at close range and small targets at medium and long distances. The model loss function consists of heatmap loss, offset loss, and size loss. Among them, the heatmap loss is a variant of Focal Loss[3], which can effectively solve the problem of imbalance between positive and negative samples, but cannot handle the imbalance between different classes. When an uncommon class appears, the class determination accuracy will be reduced. Therefore, the present invention further designs the heatmap loss and corrects it using class weights.

[0005] References

[0006] [1]Yu X, Gong Y, Jiang N, et al. Scale match for tiny person detection[C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. 2020: 1257-1265.

[0007] [2]Yu.F., Koltun.V., Multi-Scale Context Aggregation by Dilated Convolutions[C]. International Conference on Learning Representations(ICLR). 2016.

[0008] [3]Lin T Y, Goyal P, Girshick R, et al. Focal loss for dense object detection[C] / / Proceedings of the IEEE international conference on computer vision. 2017: 2980-2988. Summary of the Invention

[0009] The problem to be solved by the present invention is that in the safety anti-collision detection scenario of a rubber-tyred gantry crane, the target sizes are different at different distances, and there are differences in the existing visual detection performance. In the same visual scene, the detection performance for obstacles at different far and near distances from 0 to 20 m is poor, affecting the warning effect. The safety anti-collision of a rubber-tyred gantry crane requires a visual detection technology with excellent detection performance for both large and small targets in the detection scenario.

[0010] The technical solution of the present invention is as follows: A visual target detection method based on a hybrid convolutional residual structure constructs a visual detection model with the target detection network CenterNet, and uses the hybrid dilated convolutional residual network HDResNet as the backbone network of the visual detection model. The HDResNet is specifically as follows: Replace the 3×3 convolution in the BottleNeck of ResNet101 with a hybrid dilated convolution. The hybrid dilated convolution is three dilated convolutional kernels in parallel, and the dilation rates are 1, 2, and 5 respectively. Then use the concat module to splice the features output by the three dilated convolutional kernels;

[0011] The visual detection model is trained using an image training set. The images in the training set include targets at different visual distances. The trained model is used to perform visual detection on targets at different visual distances in the input image simultaneously.

[0012] Furthermore, residual connections are made between the Conv3 and Conv4 groups and between the Conv4 and Conv5 convolutional groups of HDResNet, and 1×1 convolutional kernels are used to adjust the dimensions and sizes.

[0013] Furthermore, when training the visual detection model, the class balance loss function BLoss is used. BLoss consists of three parts: the bias loss L off , the size loss L size and the heatmap class loss L bk . Among them, the bias loss and the size loss remain unchanged. Considering the problem of unbalanced sample classes in the training set, the heatmap class loss is designed as follows:

[0014]

[0015] where α and β are hyperparameters used to balance easy and difficult samples and positive and negative samples, N is the number of heatmaps in the image, Y xyc is the output of the heatmap localization branch. Based on the class-weighted loss algorithm of the training set statistical information, the weight w c of the class is calculated as follows:

[0016]

[0017] where, where M i represents the number of labels of the i-th class, M max and M min are the maximum number and the minimum number respectively, and γ and W are hyperparameters;

[0018] The total loss function BLoss = L bk + λ size L size + λ offset L offset , λ size = 0.1, λ offset = 1.

[0019] Furthermore, the image training set is the anti-collision image of the rubber-tyred gantry crane, which is collected by the camera deployed on the guardrail of the rubber-tyred gantry crane. The detection target is the obstacle in the rubber-tyred gantry crane scene. The distance between the detection target and the rubber-tyred gantry crane is divided into three sections, namely, 0-5m for the short distance, 5-10m for the medium distance, and 10-20m for the long distance. For the newly input anti-collision image of the rubber-tyred gantry crane, the visual detection model trained by the image training set outputs the detected visual target for the anti-collision warning of the rubber-tyred gantry crane.

[0020] The present invention also provides a visual target detection device based on a hybrid convolutional residual structure. The detection device has a computer-readable storage medium, and a computer program is configured in the computer-readable storage medium. When the computer program is executed, the above-mentioned visual target detection method is implemented.

[0021] The present invention uses CenterNet as the detection model, which can well adapt to the detection scenarios including targets at different distances, especially the anti-collision working scenario of the rubber-tyred gantry crane. Based on ResNet-101, the present invention proposes a hybrid dilated convolutional residual network HDResNet, and designs a hybrid dilated convolutional group HDC-125 from the perspective of convolutional kernels, which has the characteristic of obtaining a larger receptive field without losing information continuity; and designs a secondary residual structure, which has continuous features during the downward transmission of the network and retains a larger receptive field; at the same time, aiming at the problem of data category balance in the training set, a category balance loss function BLoss is proposed to handle the situation of category imbalance during training, reduce the influence of category imbalance in the training set, and thus increase the detection accuracy. Through experimental verification, the visual target detection method of the present invention has higher comprehensive detection accuracy at near, medium, and far distances, and has a fast detection speed. Especially in the anti-collision working scenario of the rubber-tyred gantry crane, it can effectively enhance the comprehensive detection accuracy of different targets within 0-20m of the rubber-tyred gantry crane and meet the real-time requirements of rubber-tyred gantry crane anti-collision.

[0022] The present invention uses visual detection technology to achieve the safety anti-collision of the rubber-tyred gantry crane, which has a lower cost than lidar anti-collision and has real-time performance. In the scenario of the rubber-tyred gantry crane at a distance of 0-20m, it has stronger fault tolerance than the existing visual detection technology and stronger performance in long-distance detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a schematic diagram of the anti-collision scenario of the rubber-tyred gantry crane.

[0024] Figure 2 It is a schematic diagram of the structure of the target detection network of the present invention.

[0025] Figure 3 (a) It is a schematic diagram of the existing ResNet structure.

[0026] Figure 3 (b) It is a schematic diagram of the structure of the HDResNet of the present invention.

[0027] Figure 4 It is a schematic diagram of the residual structure of the HDResNet of the present invention.

[0028] Figure 5 It is a schematic diagram of the specific application layout of the present invention in the anti-collision scenario of the rubber-tyred gantry crane. DETAILED DESCRIPTION

[0029] The present invention proposes a visual target detection method based on a hybrid convolution residual structure, which has excellent detection accuracy for objects at different distances in the detection scene. It is particularly suitable for automatic collision avoidance scenarios in tire crane work, and uses visual detection technology to ensure target detection performance at a distance of 0-20m in the tire crane work scene.

[0030] The method of the present invention uses the target detection network CenterNet to build a visual detection model, such as Figure 2 As shown in the figure, the hybrid hole convolution residual network HDResNet is used as the backbone network of the visual detection model. Specifically, the HDResNet replaces the 3×3 convolution in the BottleNeck of ResNet101 with a hybrid hole convolution, as shown in Figure 3 As shown, Figure 3 (a) is the network structure of ResNet101, with mixed hole convolution as Figure 3 As shown in (b), there are three parallel atrous convolution kernels with atrous rates of 1, 2, and 5, respectively, which ensure that the receptive field is retained as large as possible during the convolution process, and then the features output by the three atrous convolution kernels are spliced ​​using the concat module. Atrous convolution has the characteristic of retaining a larger receptive field, but it is easy to lose the continuity of information. The hybrid atrous convolution of the present invention ensures a balance between the two. The 3×3 convolution in ResNet101 is replaced with a hybrid atrous convolution. After research and analysis, it is found that the best convolution structure is to use three convolution structures with atrous rates of 1, 2, and 5, respectively. The three atrous convolutions are then concatenated using concat, and then the dimensional size is adjusted using a 1×1 convolution kernel, ensuring that the receptive field is retained as large as possible during the convolution process. For HDResNet, it is preferred to perform residual links between the Conv3 and Conv4 groups and the Conv4 and Conv5 convolution groups, that is, to directly short-circuit them, and use a 1×1 convolution kernel to adjust the dimension and size, such as Figure 4 As shown in the figure, the surface network features are further enhanced to improve the small target detection capability. After the model is built, the visual detection model is trained using an image training set. The images in the training set include targets at different visual distances. The trained model is used to perform visual detection on targets at different visual distances in the input image at the same time.

[0031] In view of the problem of class imbalance in image training sets, the present invention also designs a class balance loss function BLoss, which consists of three parts: bias loss L off , size loss L size and heat map category loss L bkAmong them, the offset loss and the size loss remain unchanged, which are the same as the common loss functions. Considering the problem of unbalanced sample categories in the training set, the heatmap loss is adjusted. The class-weighted loss algorithm based on the training set statistics, where the weight w of the class c The calculation is as follows:

[0032]

[0033] where, where M i represents the number of labels of the i-th class, M max and M min are the maximum number and the minimum number respectively, and γ, W are hyperparameters. The class weight of the most common class is 1, and the least common class is W. The heatmap loss is designed as follows:

[0034]

[0035] where α, β are hyperparameters used to balance easy and difficult samples and positive and negative samples. N is the number of heatmaps in the image. Y xyc is the output of the heatmap localization branch. Among them, the subscript c of w c represents the number of categories, and N is the number of image heatmaps or target key points. In the case of Y xyc = 1, for easily distinguishable samples, the predicted value is close to 1, while becomes smaller, ensuring that the loss result is very small, thus playing a corrective and punishing role. For difficult-to-distinguish samples, the predicted value is close to 0, increases, and it is necessary to increase the proportion of its training.

[0036] In the case of otherwise in the L bk formula, in order to prevent the predicted value from being close to 1, is used to punish the loss. And (1 - Y xyc ) β , the closer the parameter β is to the center, the smaller its value, and this weight is used to reduce the punishment intensity. If the predicted value is close to 0, Y xyc α shrinks, which can reduce the loss in this case.

[0037] The offset loss L off in BLoss is: In the whole training process, assuming that the k-th target among N targets is a certain class in the C category, the target box is expressed as Calculate the true target center point p for training, and the calculation method is For the downsampled coordinates, set them as Where R represents the downsampling multiple of 4. After quadruple downsampling, the original image target is mapped to the original image with a large error. Therefore, localoffset is additionally adopted for each center point. is the bias value output by the network and is trained with L1Loss. The bias loss is shown in the following formula, where is the predicted bias value, is the deviation value calculated during the training process.

[0038]

[0039] The size loss L in BLoss size is: The coordinate position of the center point p of the k-th target is The length and width of the target are Through L1loss training, the length is equal to the width, and the loss function is as follows. is the predicted size of the k-th target with the center point p output by the network.

[0040]

[0041] The formula for the total loss function BLoss is as follows:

[0042] BLoss = L bk + λ size L size + λ off L off

[0043] Among them, the overall loss function is the sum of the target category loss, size loss, and bias loss, and each loss has a corresponding weight. Here, the value of λ size = 0.1, λ off = 1.

[0044] The present invention is particularly applicable to the anti-collision scenario of a rubber-tyred gantry crane. In the safety anti-collision scenario of a rubber-tyred gantry crane based on visual detection, when the obstacle is at different distance positions, the rubber-tyred gantry crane processes differently. Generally, three types of operations are performed according to different distance ranges. When the obstacle is at 0 - 5m, the rubber-tyred gantry crane needs to stop urgently. When it is at 5 - 10m, the rubber-tyred gantry crane needs to decelerate and stop. When it is at 10 - 20m, the rubber-tyred gantry crane needs to give a warning. Since the size of the obstacle displayed in the image is different when the obstacle is at different distance positions from the rubber-tyred gantry crane, the obstacle at 0 - 5m is displayed larger, and the obstacle at 10 - 20m is displayed smaller. Existing visual detection has excellent performance in detecting large targets, but poor performance in detecting small targets. Therefore, the safety anti-collision of a rubber-tyred gantry crane requires a visual detection technology with excellent detection performance for both large and small targets. The implementation of the present invention will be specifically described below with the anti-collision scenario of a rubber-tyred gantry crane.

[0045] Obtaining a visual detection model trained by training an image training set, including:

[0046] Dataset collection: Data is collected using cameras deployed on the rubber-tyred gantry crane. The main target obstacles are pedestrians and container trucks, accounting for more than 80%. Other categories include toolboxes, pickup trucks, etc.

[0047] Dataset division: The dataset is divided into a training set, a validation set, and a test set, with a ratio of 7:2:1 for model training and model evaluation. To effectively evaluate the detection effect of the network at different distances, the distance between the target and the rubber-tyred gantry crane is divided into three segments, namely, a short distance of 0 - 5m, a medium distance of 5 - 10m, and a long distance of 10 - 20m.

[0048] During training and the actual detection process, it is preferred to preprocess the images: correct the distortion of the image data, adjust the image size to be consistent with the image size during training, and the input image size of the model is 640×480; then load the preprocessed image into the visual target detection model for detection.

[0049] Figure 1 Schematic diagram of the detection distance of the rubber-tyred gantry crane Figure 5 It is a schematic diagram of the rubber-tyred gantry crane safety anti-collision system. The camera is deployed at the guardrail of the rubber-tyred gantry crane, and the industrial control computer device for image processing is deployed inside the electrical room. The specific implementation of the rubber-tyred gantry crane safety anti-collision system based on this technology includes a hardware deployment stage, a system preparation stage, a detection model deployment stage, and a rubber-tyred gantry crane automatic anti-collision stage, as follows.

[0050] 1) Hardware deployment stage. An industrial control computer is arranged in the electrical room of the rubber-tyred gantry crane to process image data and execute the program of the rubber-tyred gantry crane safety anti-collision system; a camera is installed on each of the front and rear guardrails of the rubber-tyred gantry crane to collect image data, and the image data is transmitted to the industrial control computer using a POE switch; the target detected by the rubber-tyred gantry crane safety anti-collision system and the measured distance are converted into corresponding binary warning signals, which are sent to the rubber-tyred gantry crane control system through the PLC for automatic anti-collision operations. All log data and monitoring data during the system operation stage are stored in the hard disk recorder for system debugging.

[0051] 2) System preparation stage. In order to perform safety anti-collision control operations on the rubber-tyred gantry crane, the system built by the present invention needs to obtain the operating status of the rubber-tyred gantry crane. The industrial control computer receives the operating status of the rubber-tyred gantry crane sent by the rubber-tyred gantry crane control system from the PLC, mainly including the operating status of the rubber-tyred gantry crane, the operating speed of the rubber-tyred gantry crane, and the driving direction of the rubber-tyred gantry crane. The system built by the present invention needs to turn on cameras in different directions for different directions after the system starts.

[0052] 3) Detection model deployment stage, that is, deploying the visual detection model trained according to the visual target detection method of the present invention to the industrial control computer of the rubber-tyred gantry crane for real-time anti-collision detection. This stage mainly includes data preprocessing, target detection and ranging, and detection result fusion.

[0053] 3.1) Data preprocessing. After the image data and laser data are transmitted to the industrial control computer through the switch, it is first necessary to synchronize the time of the image data and the industrial control computer. The IEEE 1588 clock synchronization protocol based on Ethernet is used to synchronize the time of the front and rear cameras of the rubber-tyred gantry crane and the industrial control computer, and a timestamp is added to each frame of image data.

[0054] 3.2) Object detection and ranging. Among them, visual object detection processes the image data through the object detection model trained by the method of the present invention and outputs the detection result.

[0055] Specifically, the object measurement is carried out by the fixed-point calibration method for ranging. For the target rectangular frame detected by the visual detection model of the present invention, the distance between the real target and the rubber-tyred gantry crane is converted through the coordinate system. First, the camera is used to collect the image of the target obstacle on the driving road, and the object detection algorithm is used for detection; then the corresponding fitting rectangular frame is drawn for the detected object, and the positions of the left and right bottom corner points of the rectangular frame in the pixel coordinate system are obtained, denoted as (u 1 , v 1 ), (u 2 , v 2 ); using the pre-calibrated conversion matrix between the objective world coordinate system and the camera coordinate system of the camera, the plane coordinate points (u 1 , v 1 ), (u 2 , v 2 ) are converted into (x 1 , y 1 , z 1 ), (x 2 , y 2 , z 2 ) in the objective world three-dimensional coordinate system, where z 1 , z 2 is the required distance. Only considering the obstacles on the ground, then y 1 = 0, y 2 = 0;

[0056] 4) Automatic anti-collision stage of the rubber-tyred gantry crane. The target position and distance calculated by the application of the hybrid convolution object detection in the safety anti-collision of the rubber-tyred gantry crane are converted into corresponding deceleration and stop binary code stream signals, and are sent to the rubber-tyred gantry crane control system through the PLC. The rubber-tyred gantry crane control system performs automatic anti-collision control operations according to the early warning signal.

[0057] Based on the above scenario, in order to effectively evaluate the hybrid convolution residual network, quadratic residual, and class balance loss function BLoss of the present invention, analysis and testing are carried out in the rubber-tyred gantry crane scenario.

[0058] Table 1

[0059]

[0060] The Hybrid Dilated Convolution (HD) can effectively increase the receptive field, and its detection effect at a medium distance of 5 - 10m is better than that of the existing ResNet. As can be seen from Table 1, the standard convolution ResNet can effectively ensure the detection effect of large targets at short and medium distances. The Hybrid Dilated Convolution (HD) performs better at medium and long distances. However, as the number of stacked dilated convolution layers increases, to a certain extent, it will reduce the network learning ability. For example, HD1257 has a certain decline in detection performance compared to other Hybrid Dilated Convolutions. The HD125 of the present invention has similar detection performance to the standard convolution at short distances and the best evaluation indicators at medium and long distances.

[0061] Compared with the existing visual target detection methods, the improved solution of the present invention has stronger comprehensive performance in the target dataset of 0 - 20m for rubber-tyred gantry cranes, as shown in Table 2. FPS and mAP are two important evaluation indicators of the target detection algorithm. FPS is used to evaluate the speed of target detection, that is, the number of pictures that can be processed per second, and mAP is the accuracy of target detection.

[0062] Table 2

[0063]

[0064] As can be seen from the table, although the average detection accuracy of Faster-RCNN is good, its real-time frame rate is only 6fps, less than half of that of the one-stage detection model, which does not meet the requirements of real-time detection. The visual detection algorithm proposed based on the present invention can achieve an average detection accuracy of 76.2%, which is 3.6% higher than that of CenterNet, 3.9% higher than that of EfficientNet, and 4.4% higher than that of Scale Match. The detection frame rate is 14fps, which meets the real-time detection requirements of rubber-tyred gantry cranes. Among them, the average accuracy of HDResNet + CenterNet is 3% higher than that of ResNet + CenterNet, and the quadratic residual structure and the class balance loss function BLoss also bring a certain improvement in detection accuracy.

Claims

1. A visual object detection method based on a hybrid convolutional residual structure, characterized in that the object detection network CenterNet is used to construct a visual detection model, and the hybrid dilated convolutional residual network HDResNet is used as the backbone network of the visual detection model. The HDResNet is specifically: replacing the 3×3 convolution in the BottleNeck of ResNet101 with a hybrid dilated convolution, where the hybrid dilated convolution is three parallel dilated convolutional kernels with dilation rates of 1, 2, and 5 respectively, and then using a concat module to splice the features output by the three dilated convolutional kernels; using an image training set to train the visual detection model, where the training set images include objects at different visual distances, and the trained model is used to perform visual object detection on the input image at different visual distances simultaneously; wherein, the image training set is the anti-collision image of a rubber-tyred gantry crane, and the images are collected by a camera deployed on the guardrail of the rubber-tyred gantry crane. The detection target is the obstacle in the rubber-tyred gantry crane scene. The distance between the detection target and the rubber-tyred gantry crane is divided into three segments, namely, a short distance of 0 - 5m, a medium distance of 5 - 10m, and a long distance of 10 - 20m. The visual detection model trained by the image training set outputs the detected visual target for the newly input anti-collision image of the rubber-tyred gantry crane, which is used for the anti-collision warning of the rubber-tyred gantry crane; When training the visual detection model, the class balance loss function BLoss is used. BLoss consists of three parts, including the bias loss L off , the size loss L size and the heatmap class loss L bk . Considering the problem of class imbalance in the training set, the class weight w c is added to the heatmap class loss. L bk is designed as follows: where α and β are hyperparameters used to balance easy and hard samples and positive and negative samples, N is the number of heatmaps in the image, and Y xyc is the output of the heatmap localization branch, is its predicted value, and the class-weighted loss based on the training set statistics is used to calculate the class weight w c : Among them, M c represents the number of labels of the c-th class. M max and M min are respectively the maximum and minimum numbers of labels in the training set. γ and W are hyperparameters; The total loss function BLoss = L bk + λ size L size + λ off L off , where λ size = 0.1, λ off = 1.

2. A visual object detection method based on a hybrid convolutional residual structure according to claim 1, characterized in that residual connections are made between the Conv3 and Conv4 groups and between the Conv4 and Conv5 convolutional groups of the HDResNet, and 1×1 convolutional kernels are used to adjust the dimensions and sizes.

3. A visual object detection device based on a hybrid convolutional residual structure, characterized in that the detection device has a computer-readable storage medium, and a computer program is configured in the computer-readable storage medium. When the computer program is executed by a processor, the visual object detection method described in claim 1 or 2 is implemented.

Citation Information

Patent Citations

  • Image classification method and device

    CN111898709A

  • Weak supervision remote sensing target detection method based on hybrid hole convolution

    CN112183414A