A multi-scale object detection method and system for smoothly transmitting semantic information

By using the backbone network containing the residual block of the ReLU layer and the loss function of the balance coefficient in the object detection method, the problems of semantic information loss and class imbalance in multi-scale object detection are solved, and more accurate and robust object detection is achieved.

CN114241188BActive Publication Date: 2025-07-01CHENGDU TIANHE YICHENG TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111425402.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-26
Publication Date
2025-07-01
Estimated Expiration
2041-11-26

AI Technical Summary

Technical Problem

The existing object detection methods are insensitive to multi-scale targets, resulting in the loss of semantic information during training and the problem of class imbalance, resulting in the detector being not robust enough and the generalization is not high.

Method used

The multi-scale object detection method of the two-stage method is adopted to solve the class imbalance problem by building a backbone network containing residual blocks of the ReLU layer, smoothly transmitting semantic information, and introducing a loss function of the equilibrium coefficient.

Benefits of technology

It effectively prevents the degradation of deep networks, extracts image features in a complete and efficient manner, improves the accuracy of pixel classification, and improves the robustness and generalization capabilities of the detector.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114241188B_ABST
    Figure CN114241188B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-scale object detection method and system for smoothly transmitting semantic information, which relates to the field of computer vision technology, can improve or prevent the degradation of deep networks, extract image features completely and efficiently, and make pixel classification more accurate. The method includes: S1, constructing a residual block with a ReLU layer placed before the convolutional layer; S2, constructing a first-stage backbone network using the residual block of S1; S3, extracting features from an image using the first-stage backbone network to obtain a feature map; S4, inputting the feature map obtained in S3 into a second-stage object detection network for processing to obtain region proposals; S5, mapping the region proposals onto the feature map to obtain regions of interest, and obtaining object detection results through ROI Align processing. The second-stage object detection network integrates a cross-entropy loss function introducing a balance coefficient. The convolutional layer of the residual block performs feature image cutting and recombination operations during convolution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a multi-scale object detection method and system for smoothly transmitting semantic information. Background Art

[0002] In recent years, computer vision has flourished with the support of deep learning, and object detection has shown great practical and research value in many fields of computer vision. Many object detection methods proposed earlier were based on manually designed features. The disadvantage of these detectors is the lack of generality. Currently, more and more detection models based on convolutional neural networks (CNNs) have been proposed. By combining with CNNs, object detectors have shown powerful capabilities in handling complex tasks. There are two types of detectors: (1) one-stage methods and (2) two-stage methods. For object detectors using the two-stage method, in the first stage, the detector attempts to find a set of candidate boxes containing objects. In the second stage, these boxes will be used to introduce important semantic information for subsequent classification and regression. The two-stage method has achieved higher accuracy than the one-stage method on PASCAL VOC and MS COCO. In the one-stage method, object classification and bounding box regression are combined into one step and performed simultaneously. In 2017, a new framework detector called Region Proposal Network (RPN) was proposed in Faster R-CNN. The region proposal network will be responsible for finding the Region Of Interest (ROI) in the first stage and inputting the results into the second stage for refined regression. In the one-stage method, object detection is a regression rather than a classification task. You Look Only Once (YOLO) divides the input image into multiple parts to predict objects. Single Shot Detector (SSD) combines features from different levels in one detector for object prediction. Compared with YOLO, SSD has better performance. However, both the above one-stage method and two-stage method have some deficiencies. The main defect of the currently popular methods is the insensitivity to multi-scale objects. First, in most detection datasets, the scales of objects vary greatly. Worse still, there are some extremely large or extremely small objects. This scale variation will make it difficult for the detector to consider different scales, thus weakening the accuracy of these detectors. These methods all need to perform small-size reshaping on the input image at the beginning of training, inevitably losing some semantic information and losing more information in the subsequent convolutional process. Second, there is a class imbalance problem in existing datasets. The number of easily classifiable samples is much larger than that of difficult-to-classify samples. Therefore, the detector is not robust enough and has low generalization ability.

[0003] Therefore, it is necessary to study a multi-scale object detection method for smoothly transmitting semantic information to address the deficiencies of the prior art and solve or mitigate one or more of the above problems. Summary of the Invention

[0004] In view of this, the present invention provides a multi-scale object detection method and system for smoothly transmitting semantic information, which can improve or prevent the degradation of the deep network, extract image features completely and efficiently, and make pixel classification more accurate.

[0005] On the one hand, the present invention provides a multi-scale object detection method for smoothly transmitting semantic information, which is implemented by a two-stage method. The steps of the method include:

[0006] S1. Construct a residual block with the ReLU layer placed before the convolutional layer;

[0007] S2. Use the residual block in S1 to construct the backbone network of the first stage;

[0008] S3. Use the backbone network of the first stage constructed in S2 to extract features from the image to obtain a feature map;

[0009] S4. Input the feature map obtained in S3 into the object detection network of the second stage for processing to obtain region proposals;

[0010] S5. Map the region proposals in S4 to the feature map in S3 to obtain regions of interest, and obtain the object detection result through ROI Align processing.

[0011] In the above aspect and any possible implementation manner, a further implementation manner is provided. The object detection network of the second stage includes a region proposal network and a classifier, and a cross-entropy loss function is integrated in both the region proposal network and the classifier;

[0012] The cross-entropy loss function is:

[0013]

[0014] where α is a balance coefficient, α ∈ [0, 1];

[0015] y ∈ ±1, representing the true value;

[0016] p ∈ [0, 1], representing the probability estimation of the model for the class with the label y = 1.

[0017] In the above aspect and any possible implementation manner, a further implementation manner is provided. The work content of the convolutional layer in the residual block in step S1 includes convolution and cutting and reorganizing the feature image during the convolution process.

[0018] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The specific operation of the convolutional layer of the residual block is as follows: divide the feature map into multiple groups with the same dimension, perform SAME convolution operations, and then splice the convolution output results of all groups as the output result of this convolutional layer.

[0019] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The residual block in step S1 includes three convolutional layers.

[0020] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The first-stage backbone network in step S2 includes 164 convolutional layers, specifically including 162 convolutional layers of 54 residual blocks in four stages and two convolutional layers for adjusting the size of the picture.

[0021] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The specific content of step S4 includes: using the sliding window method to extract anchor boxes from the feature map, and then dividing it into two routes. One route predicts the probability of a positive anchor through SoftMax, and the other route preliminarily determines the preselected box through regression; then predict the probability of a positive anchor in each preselected box and output the region proposal;

[0022] The positive anchor specifically contains an object.

[0023] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The content of processing the region of interest through ROI Align (i.e., the operation method of region of interest feature alignment) in step S5 includes: further refining the position of the bounding box of the region of interest through regression, and at the same time calculating the class score of the object in the bounding box through SoftMax to further determine the object class, and obtaining the target detection result.

[0024] For the aspects and any possible implementation manners described above, a further implementation manner is provided. The residual block in step S1 includes three ReLU layers, and the ReLU layers and convolutional layers are arranged in one-to-one correspondence.

[0025] On the other hand, the present invention provides a multi-scale target detection system for smoothly transmitting semantic information, which is applicable to any of the methods described above; the system includes:

[0026] An image acquisition module for acquiring images;

[0027] A first-stage processing module for extracting features from the images acquired by the image acquisition module to obtain feature maps;

[0028] The second-stage processing module is used to process the feature map obtained by the first-stage processing module to obtain region proposals;

[0029] The combination module is used to map the region proposals of the second-stage processing module onto the feature map of the first-stage processing module to obtain regions of interest;

[0030] The ROI Align processing module is used to further process the regions of interest of the combination module to obtain object detection results.

[0031] Compared with the prior art, one of the above technical solutions has the following advantages or beneficial effects: By designing a residual block with a ReLU layer before the convolutional layer, the present invention can construct a deep network (164 layers) without losing semantic information, and can well improve or prevent the degradation of the deep network;

[0032] Another one of the above technical solutions has the following advantages or benefits: By designing a new loss function introducing a balance coefficient, the class imbalance problem can be effectively solved;

[0033] Another one of the above technical solutions has the following advantages or benefits: By cutting and reorganizing the feature map during the convolution process, the number of parameters is greatly reduced, the operation efficiency of the deep neural network is effectively improved, and the problem that the deep network is limited in use due to huge computational complexity is solved.

[0034] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned technical effects simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0036] Figure 1 It is a design diagram of an efficient residual structure provided by an embodiment of the present invention;

[0037] Figure 2 It is a design diagram of a feature channel self-learning strategy provided by an embodiment of the present invention;

[0038] Figure 3 It is a design diagram of a backbone lightweight object detection network Pre-ReNet provided by an embodiment of the present invention;

[0039] Figure 4 It is an overall architecture diagram of a two-stage object detection algorithm provided by an embodiment of the present invention;

[0040] Figure 5 It is the loss curve graph in the network prediction provided by an embodiment of the present invention;

[0041] Figure 6 It is the detection result graph combining YOLO v4 and Faster R-CNN provided by an embodiment of the present invention. Detailed implementation manners

[0042] To better understand the technical solution of the present invention, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0043] It should be clear that the described embodiments are only part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of the present invention.

[0044] To improve the accuracy and efficiency of object detection, the following two aspects should be concerned and optimized: First, in the image feature extraction stage, how the backbone extracts image features completely and efficiently, so that pixel classification is more accurate; Second, in the object localization and detection stage, how the detector accurately locates and detects the target object. No matter which aspect, it is very challenging to achieve the ideal effect, and the present invention is mainly designed and optimized for the first aspect.

[0045] The present invention proposes a new object detection backbone network called Pre-ReNet. As the core idea, the present invention proposes a Pre-ReLU residual block, aiming to prevent the degradation of deep networks and build a deep network without losing semantic information; at the same time, a new loss function is proposed to solve the extreme class imbalance problem. By using a variety of deep learning techniques and means, based on the backbone network Pre-ReLUResNet (Pre-ReNet) of the Pre-ReLU residual block and feature map cutting and recombination, the positive and negative sample imbalance in the dataset is balanced by weighting.

[0046] The main innovation points of the multi-scale object detection method for smooth transmission of semantic information of the present invention include:

[0047] (1) A new type of residual block containing the Pre-ReLU mechanism is proposed; compared with the traditional method, the activation operation is placed in front, aiming to create a "portal" for smooth transmission of semantic information.

[0048] (2) By pre-cutting and subsequent merging of the feature map, from a mathematical point of view, the number of parameters is reduced and the running efficiency of the detector is improved.

[0049] (3) A new backbone network named Pre-ReNet was created by combining Pre-ReLU and the feature map cutting and recombination mechanism.

[0050] (4) A new loss function was proposed to solve the class imbalance problem; the roles of different types of samples (such as, hard-to-classify positive samples, hard-to-classify negative samples, easy-to-classify positive samples, easy-to-classify negative samples) in the loss were analyzed, and targeted weighting was performed, thereby reducing the proportion of easy-to-classify negative samples in the loss to achieve the effect of balancing classes.

[0051] (5) The first four items were integrated to construct a new object detection neural network model named Balanced Pre-ReNet (B-PesNet) to verify its effectiveness.

[0052] ResNet solved the degradation problem of deep neural networks to a certain extent. However, after the network was deepened without limit, the network degradation problem still occurred as expected. Considering that the residual block solved the network degradation problem through skip connections that retained shallow information. After a large number of experiments and strict mathematical verification, it was found that if the ReLU layer was placed in front of the convolutional layer, a residual block using the Pre-ReLU mechanism could be obtained to improve its performance.

[0053] After the network was deepened, in the traditional method, in addition to causing the degradation problem, there were also a huge number of parameters that followed. By cutting and recombining the feature map, the number of parameters was reduced from a mathematical perspective, and the operating efficiency of the deep neural network was improved. And all convolutions used 1×1 and 3×3 convolutional kernels, and 5×5 or larger convolutional kernels were abandoned to retain more information at the convolutional level. Then, the Pre-ReLU residual block was combined with the feature map cutting and recombination to design a new backbone network named Pre-ReNet.

[0054] Balanced Loss is used to solve the imbalance problem between the foreground and background in the training stage. The class imbalance problem seriously affects the performance of Cross Entropy Loss (CE). Easy-to-classify negative samples occupy the main part in the loss and dominate the direction of gradient descent. Therefore, a reshaped loss function was proposed to reduce the proportion of easy-to-classify negative samples in the loss, focus the network on hard-to-classify negative samples in the training stage, and strengthen the learning effect targeted.

[0055] The main content of the multi-scale object detection method for smoothly transmitting semantic information of the present invention will be described in detail below:

[0056] 1. Pre-ReLU Residual Block

[0057] The standard residual block is introduced. The Pre-ReLU residual block constructed in this method is as follows:

[0058] y l = h(x l ) + F(x l , W l ) (1)

[0059] x l+1 = f(y l ) (2)

[0060] In the above formula, x l is the input of the l-th residual unit, W l refers to a set of weighted coefficients related to the l-th residual block, f represents ReLU, and h represents the skip connection: h(x l ) = x l , and F represents the residual function. As is well known, the skip connection helps to bring out the advantages of ResNet. This design can propagate more lower-layer information to the deeper layers. However, in a real training, it is found that deeper ResNet will lead to more serious degradation. In the present invention, it is assumed that f is a skip connection: x l+1 = y l , then equation (1) can be substituted into equation (2), and equation (3) is obtained:

[0061] x l+1 = x l + F(x l , W l ) (3)

[0062] For the entire network, it becomes:

[0063]

[0064] where x L represents the input of any residual block after the first unit. Then it is found that for any unit, information can be propagated from one unit to other units. To construct a "gate" that can transmit information from the shallow layer to the deep layer without loss, a ReLU layer is placed in front of the convolutional layer. As assumed, the Pre-ReLU residual block establishes a clean channel from the shallow layer to the deep layer, and more features can be retained during the convolution process. Att Figure 1 shows the structure of the efficient residual block proposed by the present invention.

[0065] 2. Feature map cutting and recombination

[0066] To reduce the number of parameters and improve the running speed, the feature map is divided into multiple groups. Att Figure 2Shows the operation process of this mechanism. First, the feature map is divided into equal dimensions at the dimension level and divided into multiple groups with the same dimension. Assume the input image size is H×W×c1, divided into g groups, and the size of each group is Corresponding to convolution kernels with a size of Perform SAME convolution operation, and the size of each convolution output is Then concatenate the g outputs, and the final output feature map size is H×W×c2. The number of parameters obtained in this process is:

[0067]

[0068] It can be seen that the number of parameters is inversely proportional to the number of groups (g). The more groups there are, the fewer parameters there are. Through such operations, the number of parameters is greatly compressed.

[0069] 3. Balanced Loss Function

[0070] To solve the extreme imbalance problem between foreground and background during training, the standard cross-entropy loss function is improved, and the cross-entropy loss function is reconstructed to reduce the weight of easily classified negative samples in training, making the network more focused on learning difficult-to-classify negative samples.

[0071] The standard cross-entropy loss function is as follows:

[0072]

[0073] Here y∈±1 represents the true value, and p∈[0,1] represents the probability estimate of the model for the class with label y = 1. For the convenience of expression, we define the probability estimate p t As follows:

[0074]

[0075] Then rewrite it as: CE(p,y) = CE(p t ) = -log(p t ).

[0076] In this method, we introduce a balance coefficient α∈[0,1], and the proposed BL is as follows:

[0077]

[0078] It is proved from theory and experiments that α can effectively balance positive and negative samples. When α < 1, 1 - α < 1, then for negative samples, the proportion they account for in the gradient will decrease, which can make both positive and negative samples dominate the gradient during training. For the convenience of expression, α is defined tAs follows:

[0079]

[0080] Therefore, BL can be rewritten as: BL(p, y) = BL(p t ) = -α t log(p t ).

[0081] 4. Construction of the Pre-ReNet Network

[0082] The backbone network designed in the present invention is called Pre-ReNet, as shown in the appendix Figure 3 . The overall structure of Pre-ReNet is shown. In Pre-ReNet, the Pre-ReLU residual block contains 3 convolutional layers. The kernel sizes of the convolutional layers are 1×1, 3×3, and 1×1 respectively. According to the size of the output image, the network can be divided into 5 parts. The first part is a 7×7 convolutional kernel, the second part contains 3 residual blocks, the third part contains 4 residual blocks, the fourth part contains 6 residual blocks, and the fifth part contains 3 residual blocks.

[0083] Generalized to the general case, the output of the first convolutional layer can be obtained as follows:

[0084]

[0085] where is the output of the (l - 1)th layer, that is, the input of this layer. f l (g) is the ReLU function, is the convolutional kernel, × represents convolutional multiplication, and the offset of this layer is

[0086] The output of the mth pooling layer is:

[0087]

[0088] where the sum of the entire feature matrix is S(G), f s (G) represents the SoftMax function, the weight coefficient is β j , and the input of this layer is and the offset is

[0089] The steps of the multi-scale object detection method for smoothly transmitting semantic information in the present invention include:

[0090] Step 1: Build a Pre-ReLU efficient residual structure, as shown in the appendix Figure 1 .

[0091] The Pre-ReLU residual block of the present invention is constructed based on the standard residual block, which was proposed in ResNet. Therefore, the present invention uses ResNet as the basic network structure. Compared with the standard residual block, in order to more smoothly transfer semantic information to deeper layers, the ReLU layer is placed before the convolutional layer in the Pre-ReLU residual block. The mathematical theory was mentioned above and will not be elaborated here. And during the convolution process, the feature map is cut to form multiple groups, and then recombined before output. The calculation process of cutting and recombining is as shown in Appendix Figure 2 as shown.

[0092] Step 2: Build the Pre-ReNet backbone network, as shown in Appendix Figure 3 as shown.

[0093] The core idea of the backbone network Pre-ReNet implemented by the present invention is the Pre-ReLU residual block, which solves the problem of network degradation in deep neural networks. Therefore, the number of layers of the backbone network Pre-ReNet designed by the present invention is set to 164 layers, divided into 4 stages. The number of Pre-ReLU residual blocks in each stage is 3, 10, 36, and 5 respectively. Since each Pre-ReLU residual block has 3 convolutional layers, there are a total of 162 layers in the 4 stages. Plus the two convolutional layers for adjusting the image size, there are 164 convolutional layers.

[0094] Step 3: Construct B-PesNet, as shown in Appendix Figure 4 as shown.

[0095] Use the backbone network Pre-ReNet built in Step 2 as the backbone network of the two-stage object detection network B-PesNet proposed by the present invention, and embed the Balanced Loss function proposed by the present invention into the multi-classifier in the final output stage, that is, use the Balanced Loss function to replace the standard cross-entropy loss function, and solve the problem of class imbalance from the perspective of the loss function.

[0096] Step 4: Process the dataset. First, download the Pascal VOC 2007 dataset on GitHub and format and install the dataset according to the instructions in the Readme document. Divide the dataset into a training set, a validation set, and a test set according to the principle of 8:1:1. Run the data augmentation algorithm provided by PyTorch on the training set to perform offline augmentation on the data, and output the augmented training set. The offline augmented set performs data augmentation operations before training, and the training set input during training is the augmented training set.

[0097] Step 5: Combine the training set and the validation set output in Step 4 (the key point here is that the validation set is not data-augmented, which can further improve the future generalization of the model from the data perspective), and use it as the training set of the present invention to input into the neural network to start training.

[0098] Specifically:

[0099] 1. Through the backbone network, extract the features in the input image and output the feature map that runs through the entire network;

[0100] 2. Divide the obtained feature map into two routes for the next step of training;

[0101] a) Input the feature map into the Region Proposal Network, and extract anchor boxes via the sliding window method. Here, it is further divided into two routes. One route uses a binary classifier to predict the probability of being a positive anchor, where a positive anchor is an anchor that contains an object. The other route uses a regressor to preliminarily determine the offset of the anchor box. Predict the probability of containing the target in each anchor box (here, the prediction only predicts whether it contains the template), and then the Region Proposal Network outputs region proposals;

[0102] b) The second route of the feature map obtained in 1 above is to input the feature map into ROI Align. By mapping the region proposals obtained in a) to the feature map, obtain the regions of interest. The regions of interest further refine the position of the bounding box through regression, and at the same time use SoftMax to predict the class score of the object in the bounding box (here, the prediction will predict what the object in the bounding box specifically is according to the classes in the training set), and then output the final result.

[0103] Thus, a trained neural network model is obtained.

[0104] Step 6: Directly input the picture to be detected into the trained neural network model. The neural network model will mark the objects detected in the picture with a square box and show the confidence level, and then save the detection result to the specified local directory, as shown in the appendix Figure 6 as shown.

[0105] Example 1:

[0106] Step 1: Image data preparation.

[0107] The datasets for training, validation, and testing in this paper are the public dataset Pascal VOC 2007. It consists of 21 classes (20 foreground and 1 background). The images in the dataset are RGB three-channel, each channel having 8-bit depth, so each image has a bit depth of 24. Considering that objects in real life are often subject to various interferences and are not as simple as the samples in the dataset. Models trained with simple datasets are difficult to apply to complex and changing real-world scenarios, that is, their generalization ability is not strong. To overcome such problems, offline data augmentation operations are introduced before training. We performed operations such as flipping, rotating, cropping, scaling, shifting, edge padding, color space conversion, noise, blurring, and random erasing without changing the relevant information of the labels to enhance the generalization ability of the model.

[0108] Step 2: Network training settings.

[0109] The programming language used in this invention is Python, version 3.6. The deep learning framework used is PyTorch, and the entire model is trained and run on a server configured with Core TM i7-6700 CPU @ 3.40GHz, GeForce GTX 1070 GPU 8GB, and 16GB RAM. The momentum stochastic gradient descent (SGD) is used as the optimization algorithm during the training process, with the momentum set to 0.9 and the weight decay set to 1e-4. The initial learning rate is set to 1e-3, and the learning rate decays to 1 / 10 of the previous batch for each batch, with the minimum learning rate set to 1e-6. 8 images are input for each batch. To explore the limit of convergence, the total number of training batches is set to 100, evenly divided into two rounds. In the first round, the parameters of Pre-ReNet are first frozen, and only a series of parameter values provided by the pre-trained model are used for training. In the second round, Pre-ReNet is unfrozen, and all network parameters are further trained.

[0110] The trained network model in this invention is saved in pth format files.

[0111] Step 3: Comparison and analysis of network parameter quantities.

[0112] Designing a network model with fewer parameters is a trend in the development of future neural network models. The fewer the parameters, the less hardware resources are required, the usage cost will be further reduced, and the usage scenarios can be further expanded, enabling deployment to more devices. In this invention, an extremely deep network is built, with the number of network layers reaching 164, which is much deeper than traditional object detection backbone networks and can extract deeper semantic information. Theoretically, the deeper the network, the larger the number of parameters and the lower the computing efficiency. However, this invention uses a large number of feature map cutting and recombination mechanisms, achieving good results in reducing the number of network parameters and realizing the lightweight of the backbone network. The number of groups (g) set in the experiment is 32. The lightweight backbone designed in this invention is experimentally compared with other excellent networks in terms of parameters, as shown in Table 1.

[0113] Table 1 Number of parameters of different backbone networks

[0114]

[0115]

[0116] As can be seen from Table 1, the network parameters designed in the present invention are far fewer than those of Darknet53, CSPDarknet53, UNet, and ResNet with various depths. This means that the backbone network designed in the present invention is smaller in size, can be deployed on a wider range of devices, and has lower requirements for devices. At the same time, if the backbone network designed in the present invention is combined with other object detection algorithms, such as YOLO v3, YOLO v4, or Faster R-CNN, the number of network parameters will be greatly reduced, and the training speed and efficiency will be greatly improved. This is exactly the purpose to be achieved by the present invention: lightweight, efficient, and compatible. Through the analysis of the data in the table, it can also be found that the number of parameters of the backbone network designed in the present invention is higher than that of the MobileNet series of networks. Objectively speaking, the efficiency is not as high as that of the current popular lightweight networks. However, the reason is that the present invention has made great efforts in the network depth in order to extract more deep semantic information features, and deepened the network depth to 164 layers, much deeper than the MobileNet series of networks. Compared with MobileNet v1, the number of network layers of the backbone network Pre-ReNet of the present invention is 5.85 times that of it, but the number of parameters is only 1.83 times that of it. Compared with MobileNet v2, the number of network layers of the backbone network Pre-ReNet of the present invention is 3.03 times that of it, and the number of parameters is 3.21 times that of it, and the efficiency is inferior to that of MobileNet v2. Compared with MobileNet v3_large, the number of network layers of the backbone network Pre-ReNet of the present invention is 2.34 times that of it, and the number of network parameters is 2.09 times that of it, which is not much different from MobileNet v3_large. It can be concluded that the backbone network designed in the present invention is excellent in terms of the number of network parameters.

[0117] Step 4: Comparison and analysis of object detection results.

[0118] The backbone network Pre-ReNet proposed in the present invention uses mechanisms such as semantic smooth transfer, feature map cutting and recombination, and foreground-background balance, retains shallow semantic information, extracts deep semantic information, reduces the number of parameters at the same time, ensures the detection accuracy, and improves the operation efficiency. In order to verify the superiority of the Pre-ReNet proposed in the present invention, based on the idea of the two-stage detection algorithm, a new object detection network framework B-PesNet using Pre-ReNet as the backbone network was constructed for experimental verification. Figure 4 shows the overall structure of B-PesNet. At the same time, in order to verify the compatibility of the Pre-ReNet proposed in the present invention, another experiment was conducted using YOLO v4. The backbone network CSPDarknet53 in YOLO v4 was replaced with the Pre-ReNet proposed in the present invention for the experiment. Figure 6The detection results of B-PesNet and Pre-ReNet combined with YOLOv4 for actual scenarios are shown. The three on the left are the detection results of B-PesNet, and the three on the right are the detection results of YOLO v4. Objectively speaking, the present invention can accurately detect the targets in the image, proving the effectiveness of the present invention.

[0119] In addition, using YOLO v3, YOLO v4, Faster R-CNN, and B-PesNet as test networks, with the Pre-ReNet designed by the present invention as the backbone and the Pascal VOC 2007 dataset as the experimental data, comparative experiments are respectively carried out with other excellent algorithms in terms of the average detection accuracy (mAP) and the operation speed (the time required to process a single image), and the results are shown in Table 2.

[0120] Table 2 Detection accuracies of different algorithms

[0121]

[0122]

[0123] It can be seen from the experimental results in the table that by comparing the changes before and after Faster R-CNN, it can be found that the accuracy has been improved, and at the same time, the efficiency of the 164-layer Pre-ReNet is 8 ms slower than that of the 50-layer ResNet50. By comparing with YOLO v3, it can be seen that the accuracy has been improved, and at the same time, the efficiency of the 164-layer Pre-ReNet is 2 ms slower than that of the 53-layer Darknet53. The experimental results before and after YOLO v4 are similar, the accuracy has been improved, and at the same time, the efficiency change is extremely small. Attached Figure 5 shows the differences in the loss decline curves before and after replacing the backbone in YOLO v4. The two curves are respectively the loss decline curve with CSPDarknet53 as the backbone and the decline curve with Pre-ReNet as the backbone. It can be clearly seen that when Pre-ReNet is used as the backbone, the loss converges faster and the final value is also lower.

[0124] The multi-scale object detection method for smoothly transmitting semantic information proposed by the present invention is mainly designed for a lightweight, highly compatible, and high-precision backbone network through theoretical design and actual verification. The backbone network Pre-ReNet of the present invention improves the detection accuracy while greatly reducing the number of parameters, making it smaller in size and easier to deploy, realizing the original intention of designing a lightweight network. Through a large number of experiments, it is proved that the method proposed by the present invention has good compatibility and strong generalization ability, can be used in current popular object detection algorithms based on convolutional neural networks, and can be directly inserted as a plugin into them to improve the efficiency and accuracy of object detection algorithms.

[0125] The above has introduced in detail a multi-scale object detection method for smoothly transmitting semantic information provided by an embodiment of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

[0126] It should also be noted that the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a commodity or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such a commodity or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the commodity or system including the said element. "Substantially" means within an acceptable error range, and those skilled in the art can solve the said technical problem within a certain error range and basically achieve the said technical effect.

[0127] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. The term "and / or" used herein is only a description of the associated relationship of the associated objects, indicating that there can be three relationships, for example, A and / or B, which can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

Claims

1. A multi-scale object detection method for smoothly transmitting semantic information, implemented by a two-stage method, characterized in that The steps of the method include: S1. Construct a residual block with a ReLU layer placed before the convolutional layer; S2. Use the residual block in S1 to construct the backbone network of the first stage; S3. Use the backbone network of the first stage constructed in S2 to extract features from the image to obtain a feature map; S4. Input the feature map obtained in S3 into the second-stage object detection network for processing to obtain region proposals; S5. Map the region proposals in S4 onto the feature map in S3 to obtain regions of interest, and obtain the object detection results through ROI Align processing; In step S1, the convolutional layer of the residual block cuts and reorganizes the feature image during convolution, specifically including: dividing the feature map into multiple groups with the same dimension; The input image size is , divided into g groups, and the size of each group is , corresponding to convolutional kernels with a size of , perform SAME convolution operation, and the size of each convolution output is , then splice the g outputs, and the size of the finally obtained output feature map is .

2. The multi-scale object detection method for smoothly transmitting semantic information according to claim 1, characterized in that The second-stage object detection network includes a region proposal network and a classifier, and a cross-entropy loss function is integrated in both the region proposal network and the classifier; The cross-entropy loss function is: ; Among them, is the balance coefficient, ; , representing the true value; , indicating the probability estimation of the model for the class labeled .

3. The multi-scale object detection method for smoothly transmitting semantic information according to claim 1, characterized in that The residual block in step S1 includes three convolutional layers.

4. The multi-scale object detection method for smoothly transmitting semantic information according to claim 3, characterized in that The backbone network of the first stage in step S2 includes 164 convolutional layers, specifically including 162 convolutional layers of 54 residual blocks in four stages and two convolutional layers for adjusting the size of the picture.

5. The multi-scale object detection method for smoothly transmitting semantic information according to claim 1, wherein The specific content of step S4 includes: using the sliding window method to extract anchor boxes from the feature map, and then dividing into two routes. One route predicts the probability of a positive anchor through SoftMax, and the other route preliminarily determines the preselected box through regression; then predict the probability of a positive anchor in each preselected box and output the region proposals; The positive anchor specifically contains an object.

6. The multi-scale object detection method for smoothly transmitting semantic information according to claim 1, wherein The content of processing the region of interest through ROI Align in step S5 includes: further refining the position of the bounding box of the region of interest through regression, and at the same time calculating the class score of the object in the bounding box through SoftMax to further determine the object class to obtain the object detection results.

7. The multi-scale object detection method for smoothly transmitting semantic information according to claim 1, wherein, The residual block in step S1 includes three ReLU layers, and the ReLU layers and convolutional layers are arranged in one-to-one correspondence.

8. A multi-scale object detection system for smoothly transmitting semantic information, characterized in that, The system can implement the steps of the method as described in any one of claims 1-7; the system includes: An image acquisition module for acquiring images; A first-stage processing module for extracting features from the images acquired by the image acquisition module to obtain a feature map; A second-stage processing module for processing the feature map obtained by the first-stage processing module to obtain region proposals; A combination module for mapping the region proposals of the second-stage processing module onto the feature map of the first-stage processing module to obtain regions of interest; An ROI Align processing module for further processing the regions of interest of the combination module to obtain object detection results.