Target detection method, device and equipment
By introducing multi-branch fusion modules and deformable convolutions into the detection model, the missed detection and misdetection problems of small object detection in low-light environments are solved, and higher detection accuracy and performance are achieved.
Patent Information
- Application Number
- CN202510205776.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
AI Technical Summary
In low-light environments, small target detection is prone to missed or misdetected, and the prior art is difficult to effectively identify the edge and texture characteristics of small targets.
Using an improved detection model based on YOLOv8, this model introduces a multi-branch fusion module in the feature fusion network, captures information at different scales through multiple parallel branches, and uses deformable convolution in the feature extraction network to enhance the extraction ability of small-objective features.
It significantly reduces the risk of missed and missed detection of small target detection in low-light environments, improves the accuracy and performance of detection, and can better identify and detect various scale changes of small targets.
Smart Images

Figure CN120125839A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular, to a method, apparatus, and device for object detection. Background Art
[0002] With the rapid development of the industrialization process, the demand for detection in technical fields such as security monitoring, medical image analysis, intelligent transportation, and power systems is increasing day by day. Generally speaking, the detection of specific objects is based on the images collected in the scene. However, in some scenes, the lighting conditions are poor, and it is more difficult to identify small objects. The risk of missed detection or false detection increases for some existing technologies that use models to detect small objects. In addition, due to the small area occupied by some small objects in the image and the insignificant edge and texture features, the detection difficulty is further increased. Overcoming the problem of small object detection in low-light environments and reducing the risk of missed detection and false detection are major challenges in object detection. Solving this problem has great significance in various industrial scenarios. Summary of the Invention
[0003] Based on the above problems, the present application provides a method, apparatus, and device for object detection, aiming to reduce the risk of missed detection and false detection of small object detection in low-light environments and achieve object detection more accurately.
[0004] The embodiments of the present application disclose the following technical solutions:
[0005] The first aspect of the present application provides a method for object detection, and the method includes:
[0006] Obtain an initial captured image of an object of a target type in a detection scene; the first additional object of the object is the detection target, and the size of the first additional object is smaller than the size of the main structure of the object;
[0007] Perform image preprocessing on the initial captured image to construct an input image for a detection model; the network structure of the detection model is a network structure improved based on YOLOv8, which includes a feature extraction network, a feature fusion network, and a prediction output network; a multi-branch fusion module is provided in the feature fusion network, and the multi-branch fusion module includes multiple parallel branches, and the multiple parallel branches are respectively used to capture information of different scales in the input feature map of the multi-branch fusion module;
[0008] After the input image is input into the detection model, image analysis is performed through the detection model to obtain an image object detection result output by the detection model; the image object detection result is used to indicate whether the object has the first additional object or not.
[0009] In an alternative implementation, during the process of performing image analysis through the detection model, the processing steps of the multi-branch fusion module include:
[0010] Perform a first convolution operation on the input feature map to obtain a first convolution result;
[0011] Divide the information in the first convolution result into first part information and second part information according to channels;
[0012] Through the multiple parallel branches, respectively perform feature extraction of different scales on the first part information to obtain multiple feature extraction results; the multiple feature extraction results correspond one-to-one to the multiple parallel branches;
[0013] Perform channel information superposition on the multiple feature extraction results and the second part information to obtain a superposition result;
[0014] Perform a second convolution operation on the superposition result to obtain a second convolution result as the output feature map of the multi-branch fusion module.
[0015] In an alternative implementation, the dividing the information in the first convolution result into first part information and second part information according to channels includes:
[0016] Determine the total number of channels of the first convolution result, and determine the multi-scale perception proportionality coefficient of the multi-branch fusion module; the multi-scale perception proportionality coefficient is a positive number less than 1;
[0017] Based on the total number of channels, extract the information of the corresponding proportion of channels from the first convolution result as the first part information according to the multi-scale perception proportionality coefficient, and use the remaining information in the first convolution result as the second part information.
[0018] In an alternative implementation, the multiple parallel branches include a first branch, a second branch, and a third branch, where the receptive field of the first branch is smaller than the receptive field of the second branch, and the receptive field of the third branch is larger than the receptive field of the second branch;
[0019] The first branch includes a standard convolutional layer with a convolutional kernel of a first size;
[0020] The second branch includes a first convolutional layer, a second convolutional layer, and a third convolutional layer. Among them, the first convolutional layer is used to convolve the horizontal features of the image, and the third convolutional layer is used to convolve the vertical features of the image; the second convolutional layer is located between the first convolutional layer and the third convolutional layer, and the second convolutional layer is a standard convolutional layer with a convolutional kernel of a second size, and the second size is greater than the first size; the third branch includes a module based on a spatial attention mechanism and / or a channel attention mechanism.
[0021] In an alternative implementation, the third branch specifically includes a dual-domain channel attention module and a frequency-based spatial attention module; the dual-domain channel attention module is located at the front end of the frequency-based spatial attention module, and the output of the dual-domain channel attention module is used as the input of the frequency-based spatial attention module.
[0022] In an alternative implementation, in both the feature extraction network and the feature fusion network, a plurality of DC2f modules are provided; the DC2f module is a deformable C2f module with deformable convolution; the positions of the DC2f modules in the network structure of the detection model are the same as the positions of the C2f modules in the network structure of YOLOv8.
[0023] In an alternative implementation, the image preprocessing of the initial acquired image to construct the input image of the detection model includes:
[0024] Convert the initial acquired image from the RGB color space to the Lab color space to obtain a first converted image;
[0025] For the first converted image, determine the shadow area based on the L channel, a channel, and b channel;
[0026] Perform channel value correction on the shadow area in the Lab color space;
[0027] Convert the image with the corrected shadow area from the Lab color space back to the RGB color space to obtain a second converted image as the input image of the constructed detection model.
[0028] In an alternative implementation, the determining the shadow area based on the L channel, a channel, and b channel for the first converted image includes:
[0029] Determine the L-channel shadow screening threshold of the first converted image according to the L-channel values of the pixels in the first converted image, and determine the b-channel shadow screening threshold of the first converted image according to the b-channel values of the pixels in the first converted image;
[0030] If the color information of the first converted image meets the first condition, determine the shadow area in the first converted image according to the L-channel shadow screening threshold; if the color information of the first converted image does not meet the first condition, determine the shadow area in the first converted image according to the L-channel shadow screening threshold and the b-channel shadow screening threshold; the first condition is: the sum of the average values of the a-channel and the b-channel of each pixel in the first converted image is less than or equal to the maximum value of the color channels of the first converted image; the maximum value of the color channels is the maximum value among the a-channel values of each pixel and the b-channel values of each pixel.
[0031] In an alternative implementation, the determining the L-channel shadow screening threshold of the first converted image according to the L-channel values of each pixel in the first converted image includes:
[0032] Determine the average value and the standard deviation of the L-channel of each pixel in the first converted image;
[0033] Multiply a first proportionality coefficient by the standard deviation of the L-channel to obtain a first product; the first proportionality coefficient is a positive number less than 1;
[0034] Calculate the difference between the average value of the L-channel and the first product as the L-channel shadow screening threshold of the first converted image;
[0035] The determining the b-channel shadow screening threshold of the first converted image according to the b-channel values of each pixel in the first converted image includes:
[0036] Determine the average value and the standard deviation of the b-channel of each pixel in the first converted image;
[0037] Multiply a second proportionality coefficient by the standard deviation of the b-channel to obtain a second product; the second proportionality coefficient is a positive number less than 1;
[0038] Calculate the difference between the average value of the b-channel and the second product as the b-channel shadow screening threshold of the first converted image.
[0039] In an alternative implementation, the performing channel value correction on the shadow area in the Lab color space includes:
[0040] Obtain the average LAB value of the shadow area and obtain the average LAB value of the non-shadow area in the first converted image;
[0041] Calculate the quotient of the average LAB value of the non-shadow area and the average LAB value of the shadow area as the correction proportionality factor of the first converted image;
[0042] Use the calibration scale factor to calibrate the LAB values of each pixel in the shadow area, and obtain the image after calibrating the shadow area.
[0043] In an alternative implementation, the conversion of the image after calibrating the shadow area from the Lab color space back to the RGB color space includes:
[0044] Smooth the edges of the shadow area in the image after calibrating the shadow area through a median filter to obtain an image after edge smoothing processing;
[0045] Convert the image after edge smoothing processing from the Lab color space back to the RGB color space.
[0046] A second aspect of the present application provides a target detection device, which includes:
[0047] An image acquisition module, configured to acquire an initial captured image of an object of a target type in a detection scene; the first additional object of the object is a detection target, and the size of the first additional object is smaller than the size of the main structure of the object;
[0048] An image preprocessing module, configured to perform image preprocessing on the initial captured image to construct an input image of a detection model; the network structure of the detection model is a network structure improved based on YOLOv8, including a feature extraction network, a feature fusion network, and a prediction output network; a multi-branch fusion module is set in the feature fusion network, and the multi-branch fusion module includes multiple parallel branches, and the multiple parallel branches are respectively used to capture information of different scales in the input feature map of the multi-branch fusion module;
[0049] A target detection module, configured to, after inputting the input image into the detection model, perform image analysis through the detection model to obtain an image target detection result output by the detection model; the image target detection result is used to indicate whether the object has the first additional object.
[0050] A second aspect of the present application provides a target detection device, which includes: a memory and a processor;
[0051] The memory is used to store a computer program;
[0052] The processor is configured to run the computer program, and when the computer program runs, it executes the steps of the method described in any implementation manner in the first aspect.
[0053] Compared with the prior art, the present application has the following beneficial effects:
[0054] Based on the initial captured image of an object of the target type in the detection scenario, this application performs image preprocessing to construct the input image of the detection model. The first additional object of the object of the target type is the detection target of the detection model. By capturing an image containing the object of the target type and performing image analysis with the detection model, it is detected whether the object in the image has the first additional object, that is, it is determined whether the first additional object exists. This detection model is improved based on the network structure of YOLOv8, and a multi-branch fusion module is set in the feature fusion network of the model. The multi-branch fusion module has multiple parallel branches, which are respectively used to capture information of different scales in the input feature map of the multi-scale fusion module. Since the multi-branch fusion module in the detection model can provide multi-scale receptive fields, the application of the multi-branch fusion module extracts image features from multiple scales, which can cover various scale changes of small targets in the captured image. In this way, the ability of the detection model to extract and recognize small target features is greatly enhanced, and the detection performance of the model for small targets in low-light environments is improved. The technical solution of this application also has high application prospects in low-light environments. After experimental verification, this solution can reduce the risks of missed detection and false detection of small target detection in low-light environments and more accurately achieve target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 It is a flowchart of a method for target detection provided by an embodiment of the present application;
[0057] Figure 2 It is a schematic diagram of the network structure of a detection model provided by an embodiment of the present application;
[0058] Figure 3 It is a schematic diagram of the structure of a multi-branch fusion module provided by an embodiment of the present application;
[0059] Figure 4 It is a schematic diagram of the structure of DCAM of the third branch in the multi-branch fusion module of the detection model;
[0060] Figure 5 It is a schematic diagram of the structure of FSAM of the third branch in the multi-branch fusion module of the detection model;
[0061] Figure 6 It is a schematic diagram of the structure of DC2f provided by an embodiment of the present application;
[0062] Figure 7 Schematic diagram of the network structure of another detection model provided by an embodiment of the present application;
[0063] Figure 8 Schematic diagram of the training and application of a detection model provided by an embodiment of the present application;
[0064] Figure 9 Schematic diagram of the process for preprocessing an image provided by an embodiment of the present application;
[0065] Figure 10 Schematic diagram of positive and negative samples in the training dataset;
[0066] Figure 11 Schematic diagram of the effect comparison of the first group of seal detections provided by an embodiment of the present application;
[0067] Figure 12 Schematic diagram of the effect comparison of the second group of seal detections provided by an embodiment of the present application;
[0068] Figure 13 Schematic diagram of the effect comparison of the third group of seal detections provided by an embodiment of the present application;
[0069] Figure 14 Schematic diagram of the effect comparison of the fourth group of seal detections provided by an embodiment of the present application;
[0070] Figure 15 Schematic diagram of the effect comparison of the fifth group of seal detections provided by an embodiment of the present application;
[0071] Figure 16 Schematic diagram of the effect comparison of the sixth group of seal detections provided by an embodiment of the present application;
[0072] Figure 17 Schematic diagram of the structure of a target detection device provided by an embodiment of the present application. Detailed implementation manners
[0073] As described above, the demand for object detection in various industrial scenarios is increasing. Below, taking the technical scenario of the power system as an example, the demand for object detection will be introduced. With the acceleration of the urban modernization process in China, the level of intelligence and automation of the power system has been continuously improved. In order to meet the growing urban electricity demand, the construction of smart grids has gradually become the core direction of the development of the power industry. As an important part of the end of the power system, the electricity meter not only undertakes the task of measuring the electricity consumption of users, but also has important significance for the operation monitoring and management of the power system. The safety and reliability of the electricity meter are directly related to the fairness of power transactions and the vital interests of users. Therefore, it is crucial to ensure the integrity and safety of the electricity meter. The electricity meter seal is mainly used to prevent illegal disassembly or tampering with the meter, ensuring that the internal parameter settings of the electricity meter are not maliciously changed, thus maintaining the fairness of power transactions. However, in actual applications, the electricity meter seal faces many challenges: on the one hand, due to the installation location usually being outdoors or in a semi-closed environment, it is easily affected by natural factors such as weather changes, humidity, and temperature fluctuations; on the other hand, the seal may be damaged by humans or aged and fall off due to long-term use, which may lead to the loss or failure of the seal, thereby affecting the safety and accuracy of the electricity meter.
[0074] Traditional seal inspection mainly relies on manual patrols. This method is not only inefficient and costly, but also difficult to ensure the comprehensiveness and timeliness of the inspection. In recent years, with the development of computer vision technology and deep learning algorithms, image-based object detection technology has begun to be applied in the field of industrial defect detection and achieved remarkable results. However, the performance of some detection models is easily affected significantly in low-light environments. This is mainly because the algorithms of the detection models are sensitive to lighting conditions. In a low-light and low-illuminance environment, the features in the image may not be obvious, making it difficult for the model to accurately capture the details of small target objects, thus increasing the risk of missed detection or false detection.
[0075] In view of this, the inventors in this application propose a method, device, and equipment for object detection, which can automatically achieve small target detection, perform outstandingly even in low-light environments, reduce the risk of missed detection and false detection of small target detection in low-light environments, and more accurately achieve object detection.
[0076] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0077] See Figure 1, This figure is a flowchart of a method for object detection provided by an embodiment of the present application. As Figure 1 shown, the method includes:
[0078] S101. Obtain an initial captured image of an object of a target type in a detection scene.
[0079] In the embodiment of the present application, when collecting images, images of objects of the same target type are collected. Taking the technical scenario of the power system as an example, if it is necessary to detect the lead seal of a watt-hour meter, an image of the watt-hour meter needs to be collected. In this scenario, the watt-hour meter is an object of the target type. The object of the target type has one or more targets to be detected. Taking the watt-hour meter as an example, it has multiple additional objects such as a display screen, a barcode label, and a lead seal of the watt-hour meter. In the detection scene of the lead seal of the watt-hour meter, the lead seal of the watt-hour meter is the detection target on the watt-hour meter. For the convenience of description, the embodiment of the present application proposes that the object of the target type should have a first additional object, and this first additional object is the detection target. The purpose of implementing the solution is to detect the presence or absence of the first additional object. In practical applications, the size of the first additional object is smaller than the size of the main structure of the object. Therefore, when the main structure of the object of the target type appears in the captured image, generally speaking, the size of the first additional object in the image is small, resulting in difficulties in identifying and detecting the target.
[0080] In the embodiment of the present application, it is proposed to use a pre-trained detection model to detect whether the first additional object is included in the image, and then determine whether the first additional object is missing and whether the object of the target type has safety and accuracy. The scenario of detecting the lead seal of the watt-hour meter mentioned above can be referred to. Generally speaking, there may be situations where the initial captured image does not match the input image required by the model in terms of image quality, image size, etc. Therefore, directly inputting the initial captured image into the detection model is likely to produce unsatisfactory detection results. In order to promote effective object detection, in the embodiment of the present application, the input image of the detection model is constructed based on the initial captured image. See step S102 for details.
[0081] S102. Perform image preprocessing on the initial captured image to construct the input image of the detection model.
[0082] In a possible implementation manner, the image preprocessing can be the processing of image quality, the processing of image size, or both. The purpose of performing preprocessing on the initial captured image is to convert the initial captured image into a form that meets the input requirements of the detection model and to enable the detection model to achieve a better detection effect based on the preprocessed image compared to directly processing the initial captured image. Since there are various ways of image preprocessing, the specific implementation of step S102 in the present application is not limited.
[0083] To facilitate the understanding of the operation of the detection model in S103 below, the network structure of the detection model will be described below. In the embodiments of the present application, the network structure of the detection model is a network structure improved based on YOLOv8. That is to say, the basic framework of the network structure of the detection model is the framework of YOLOv8. However, on the basis of the network structure of YOLOv8, in order to meet the requirements of small target detection in low-light environments, improvements have been made.
[0084] The low-light environment mentioned here is a lighting environment with insufficient light, which is the opposite of the condition of sufficient light in which the algorithm or model can clearly identify the detected object. A common definition is that the low-light environment refers to the light condition where the ambient illuminance is below EV5 or 80 lux. Another common definition is that the low-light environment refers to the light environment where the illuminance is below 50 lumens. The low-light environment mentioned in the technical solution of the present application refers to the light condition where the ambient light is insufficient, so that the detection accuracy and precision are insufficient due to the high detection difficulty when the existing detection model detects small targets.
[0085] Figure 2 It is a schematic diagram of the network structure of a detection model provided by an embodiment of the present application. As Figure 2 shown, the detection model includes a feature extraction network 21, a feature fusion network 22, and a prediction output network 23. Among them, the feature extraction network 21 can also be called the Backbone backbone network, the feature fusion network 22 can also be called the Neck neck, and the prediction output network 23 can also be called the Head head.
[0086] The feature extraction network 21 is responsible for feature extraction, and it uses a series of convolutional and deconvolutional layers. And residual connections and Bottleneck Blocks are also used.
[0087] The feature fusion network 22 enhances the feature representation ability by fusing the feature maps from different stages of the feature extraction network 21.
[0088] The prediction output network 23 is responsible for target detection and classification tasks, and can generate and output detection results.
[0089] It should be particularly noted that in the detection model mentioned in the embodiments of the present application, different from the network structure of the current existing YOLOv8 model, a multi-branch fusion module 220 is set in the feature fusion module 22 of the detection model, as Figure 2As shown in the figure. The multi-branch fusion module 220 set in the feature fusion module 22 provides an important auxiliary function for the detection model to recognize the first additional object with a smaller size in the image. Based on the structure and working mechanism of the multi-branch fusion module 220, the perception ability of the detection model in small object detection is improved, the vulnerability in feature perception is reduced, and thus the accuracy of small object detection is improved.
[0090] The multi-branch fusion module 220 has multiple parallel branches. Taking Figure 2 as an example, three parallel branches are shown, namely the first branch 221, the second branch 222, and the third branch 223. It should be noted that Figure 2 only three parallel branches are used as examples to show the structure of the multi-branch fusion module. In actual applications, according to the requirements for the perception ability of features of different sizes, more or fewer parallel branches can also be set. Therefore, the number of parallel branches in the multi-branch fusion module in the detection model is not limited here.
[0091] S103. After the input image is input into the detection model, image analysis is performed through the detection model to obtain the image object detection result output by the detection model.
[0092] After the input image is input into the detection model, it is processed by each network in the detection model in turn. First, it is processed by the feature extraction network, then by the feature fusion network, and finally by the prediction output network.
[0093] The feature extraction network adjusts the size and shape of the convolution kernel in different regions of the input image to better capture the detailed features of the image. Taking the scene of electric energy meter seal detection as an example, the image of the electric energy meter is input into the feature extraction network. Under the processing of the feature extraction network, the key features such as the electric energy meter, background, and seal in the image are extracted for further feature mining, feature processing, and classification by subsequent modules.
[0094] In the feature fusion network, multiple parallel branches in the multi-branch fusion module are respectively used to capture information of different scales in the input feature map of the multi-branch fusion module. That is to say, the information scales captured by different branches are different, which makes the receptive fields of multiple parallel branches different and increases the identification ability of the detection model for the detected target. Taking the detection of the lead seal of an electricity meter in a low-light environment as an example, in a low-light environment, the lead seal of the electricity meter is small in size, so its color and texture are likely to be less prominent due to weak light, resulting in increased detection difficulty. Using existing models, such as the traditional YOLOv8 model, it is also easy to have problems of missed detection or false detection. For example, there are actually 8 electricity meters in the scene, and there is actually a lead seal on each electricity meter, but in the image recognition and detection stage, the model identifies 8 lead seals as 6 lead seals, because the features of 2 lead seals are not prominent enough and the model fails to identify them. Or, for another example, the lead seal of the electricity meter in the scene no longer exists, but due to the shadow effect of other small-size interfering objects under light projection, it is misidentified as the lead seal of the electricity meter by the model, resulting in a wrong detection conclusion. In the embodiment of the present application, the detection model can perform small target recognition on the image of the object of the target type collected in a low-light environment, and the multi-branch fusion module in the feature fusion network of the model can enhance the extraction and recognition ability of small target features by providing multi-scale receptive fields. Even if the first additional object has size variations in different images, or the low-light environment increases the detection difficulty (such as increasing the difficulty of recognizing the edges and textures of small targets), accurate detection results can still be obtained using this detection model.
[0095] Finally, the prediction output network outputs the image target detection result to indicate whether the object has the first additional object or not. For the case where the image target detection result indicates that the object of the target type has the first additional object, the prediction output network can also output a probability value. Through the probability value, numerically represents the probability that the target recognized by the model prediction is the first additional object.
[0096] To facilitate understanding of the working mode of the multi-branch fusion module, the following Figure 3 exemplarily shows the structure of a multi-branch fusion module and Figure 3 illustrates its processing steps. Figure 3 FIG. is a schematic structural diagram of a multi-branch fusion module provided by an embodiment of the present application. As shown in the figure, the multi-branch fusion module includes two convolutional blocks, one before and one after, each used to perform a convolutional operation. For the convenience of distinction and description, the convolutional operation performed by the convolutional block with the data stream in front is called the first convolutional operation, and the convolutional operation performed by the convolutional block with the data stream behind is called the second convolutional operation. In Figure 3 this, both of these convolutional blocks are represented by rectangular blocks marked with the word "Conv". InFigure 3 In the multi-branch fusion module 220 shown, it includes a first branch 221, a second branch 222, and a third branch 223.
[0097] In Figure 3 In the structure shown, for the multi-branch fusion module 220, after the input feature map enters the module, it will first pass through a convolutional block to perform a first convolution operation, and the execution result of this operation is called the first convolution result. In the first convolution result, a part serves as the first part of information and enters the first branch 221, the second branch 222, and the third branch 223; another part serves as the second part of information and directly flows to the processing ends of the three branches without processing to be concatenated and superimposed with the processing results of the three branches 221-223. For ease of explanation, the data stream that transmits the second part of information alone is called the self-branch, and the three branches 221-223 are collectively called the multi-scale branches. The self-branch directly transmits a part of the input information, that is, the second part of information, to the output end, alleviates the problems of gradient disappearance and gradient explosion, accelerates model convergence, retains important information in the original input, and plays a role in information transmission and fusion.
[0098] As Figure 3 shown, in an alternative implementation, during the process of image analysis through a detection model, the processing steps of the multi-branch fusion module include: performing a first convolution operation on the input feature map to obtain a first convolution result; dividing the information in the first convolution result into a first part of information and a second part of information according to channels; through multiple parallel branches, respectively performing feature extraction of different scales on the first part of information to obtain multiple feature extraction results; the multiple feature extraction results correspond to the multiple parallel branches one by one; superimposing the channel information of the multiple feature extraction results and the second part of information to obtain an information superposition result; performing a second convolution operation on the information superposition result to obtain a second convolution result as the output feature map of the multi-branch fusion module.
[0099] As mentioned before, the information in the first convolution result is divided into a first part of information and a second part of information according to channels. Among them, the first part of information is given to the multi-scale branches and provided to the first branch 221, the second branch 222, and the third branch 223 in parallel without distinction, and the second part of information is given to the self-branch. The following introduces an example implementation of dividing the first part of information and the second part of information according to channels.
[0100] First, for the first convolution result obtained through the first convolution operation, determine the total number of channels of the first convolution result and determine the multi-scale perception ratio coefficient of the multi-branch fusion module. For example, the total number of channels is denoted as C, and the multi-scale perception ratio coefficient is denoted as e, where the multi-scale perception ratio coefficient e is a positive number less than 1. Next, based on the total number of channels C and the multi-scale perception ratio coefficient e, extract the information of the corresponding proportion of channels from the first convolution result as the first part of the information, and use the remaining information in the first convolution result as the second part of the information. Specifically, take the information of C*e channels as the first part of the information, and take the information of the remaining C*(1 - e) channels as the second part of the information. As an example, e = 0.25, that is, the information of one-fourth of the channels enters the multi-scale branch, and then each of the multiple parallel branches plays its own role to achieve multi-scale perception; the information of three-fourths of the channels enters the self-branch.
[0101] In an alternative implementation, the multiple parallel branches include Figure 3 the first branch 221, the second branch 222, and the third branch 223 shown in the figure. Among them, the receptive field of the first branch 221 is smaller than the receptive field of the second branch 222, and the receptive field of the third branch 223 is larger than the receptive field of the second branch 222. That is to say, from the first branch 221 to the second branch 222 and then to the third branch 223, the size of the receptive field increases step by step.
[0102] In an example structure, such as Figure 3As shown in the figure, the first branch 221 in the multi-branch fusion module 220 includes a standard convolutional layer A1 with a convolutional kernel of the first size. The second branch 222 includes a first convolutional layer B1, a second convolutional layer B2, and a third convolutional layer B3. Among them, the first convolutional layer B1 is used to convolve the features in the horizontal direction of the image, and the third convolutional layer B3 is used to convolve the features in the vertical direction of the image; the second convolutional layer B2 is located between the first convolutional layer B1 and the third convolutional layer B3, and the second convolutional layer B2 is a standard convolutional layer with a convolutional kernel of the second size. The second size is greater than the first size. This means that the second branch 222 has a larger receptive field compared to the first branch 221. As an example, the convolutional kernel of the first size is a 1×1 convolutional kernel, and the convolutional kernel of the second size is a K×K convolutional kernel. K is an integer greater than 1. The first convolutional layer B1 can be of the size 1×K, and the third convolutional layer B3 can be of the size K×1. In the example, K = 63, that is, the first convolutional layer B1 is 1×63, the second convolutional layer B2 is 63×63, and the third convolutional layer B3 is 63×1. The second branch 222 extracts rich local structures and context information through depth convolutions with different shapes and extremely large kernel sizes, can capture medium-scale features, and uses a standard convolutional layer with a medium receptive field size to ensure the balance between global context and local details. The first branch 221 extracts key features from a specific local area, uses 1×1 depth convolution to supplement the degraded local information of small sizes and performs local signal modulation, focuses on the local detail features of the image, and is dedicated to extracting the fine-grained features crucial for small target detection.
[0103] In the multi-branch fusion module, the third branch includes a module based on a spatial attention mechanism and / or a channel attention mechanism. In one example, the third branch includes a module based on a spatial attention mechanism; in another example, the third branch includes a module based on a channel attention mechanism; in yet another example, the third branch includes a module based on both a spatial attention mechanism and a channel attention mechanism.
[0104] Spatial Attention Mechanism (SAM) is a mechanism that focuses on the importance assignment of the spatial dimension of the feature map. The spatial attention mechanism weights specific spatial positions in the feature map to highlight the regions that contribute most to the task, suppress irrelevant or redundant regions, and improve the performance of the model. Channel Attention Mechanism (CAM) is a mechanism that focuses on the importance assignment of the feature map channels in Convolutional Neural Networks (CNN). The main purpose of the channel attention mechanism is to emphasize the channels that contribute most to the task by assigning different weights to each channel, suppress irrelevant or redundant channels, and thus improve the performance of the model. In the third branch, the use of modules based on the spatial attention mechanism and / or the channel attention mechanism can expand the receptive field, making the receptive field of the third branch larger than that of the second branch, thereby capturing global context information without significantly increasing the number of parameters.
[0105] In Figure 3 the example, as an optional implementation, the third branch 223 specifically includes a dual-domain channel attention module C1 and a frequency-based spatial attention module C2. Among them, the dual-domain channel attention module C1 is located at the front end of the frequency-based spatial attention module C2, and the output of the dual-domain channel attention module C1 serves as the input of the frequency-based spatial attention module C2.
[0106] The Dual-domain Channel Attention Module (DCAM) can combine spatial information and channel information to dynamically reallocate the focus of the model, making it more focused on key features. Thereby improving the detection performance of the detection model and enhancing its detection accuracy. DCAM achieves global perception through dual-domain channel attention, captures the overall information of the image and the associations between elements, and helps the model understand the content of the image from a macroscopic level. Figure 4 It is a schematic diagram of the structure of DCAM in the multi-branch fusion module of the detection model. This figure exemplarily shows the structure of DCAM, such as Figure 4 , DCAM is divided into two parts: FCA and SCA. Among them, FCA represents a module that enhances the attention in the frequency domain on the feature channels, and SCA represents a module that enhances the attention in the spectral domain on the feature channels. In FCA, FFT represents the Fast Fourier Transform, IFFT represents the Inverse Fast Fourier Transform. GAP represents global average pooling. As Figure 4 shown, FCA is located at the front end of SCA, and the output of FCA serves as the input of SCA.
[0107] The frequency-gated spatial attention module (FSAM) achieves global perception through the frequency gating mechanism, captures the overall information of the image and the relationship between elements, and helps the model understand the image content from a macro level. FSAM enhances the attention of the frequency domain in the feature space. Figure 5 The schematic diagram of the structure of the FSAM of the third branch in the multi-branch fusion module of the detection model is shown in FIG. In the figure, FFT stands for Fast Fourier Transform, and IFFT stands for Inverse Fast Fourier Transform.
[0108] In actual applications, when performing image acquisition for a target type of object, it is easy for the size of the first additional object to be different in different images. The reasons for this may be: the orientation or posture of the first additional object in different images is different, it is blocked by shadows, or the distance between the image acquisition device and the object is inconsistent. In the technical solution of the present application, a multi-channel fusion module is used in the detection model, and multiple parallel branches are used to extract information of different scales, which can well cover various scale changes of small targets, thereby significantly improving the detection performance and improving the applicability of the model to detecting different images of target size changes.
[0109] In order to suppress the computational overhead caused by the convolution of large convolution kernels in the multi-branch fusion module, in an embodiment of the present application, the multi-branch fusion module can be deployed at the convolution layer with the smallest feature map size. In addition, the receptive field is large, so that the third branch can effectively understand the image content. In this way, the multi-branch fusion module can not only give full play to its powerful feature extraction capabilities under the premise of controlling computational overhead, but also improve the understanding and analysis capabilities of the image with the help of a larger receptive field, provide high-quality feature expressions for subsequent model tasks, and thus improve the performance of the entire model, which is suitable for real-time processing of targets.
[0110] In an optional implementation, the multi-branch fusion module can use an Omni-Kernel Module (OKM). OKM was originally designed for image super-resolution tasks. In the embodiments of the present application, the inventor applies OKM to the YOLOv8 model to improve the model structure and improve the performance of the detection model for small target detection in low-light environments. In the specific processing flow of OKM, after each convolution completes feature extraction, a batch normalization layer (BN) and an activation function (Sigmoid Linear Unit, SiLU) are used.
[0111] The example implementation of the multi-branch fusion module in the detection model was introduced above. As mentioned above, the use of the multi-branch fusion module is a major improvement in the network structure of the model for small target detection in a low-light environment in the embodiment of the present application. In addition, the present application also proposes another improvement to the network structure of YOLOv8: using the DC2f module to replace the original C2f modules in the YOLOv8 network structure.
[0112] The C2f (Convolutional Block with Two Feature maps) module is an important component of the backbone network and neck of YOLOv8, responsible for extracting feature maps. In the embodiment of the present application, the inventor introduces deformable convolution (Deformable Convolution) into the C2f module to form a deformable C2f module (Deformable ConvolutionalBlock with Two Feature maps, DC2f), which can significantly improve the model's ability to capture targets of different shapes, especially in the detection of small targets in low-light environments.
[0113] The traditional convolution operation is a method of feature extraction based on a fixed spatial position. After convolution, any position p on the resulting feature map y 0 The values are:
[0114]
[0115] Among them, x is the input feature map, R defines a convolution kernel of a certain size, w is the convolution kernel parameter, and p 0 is p on the input-output feature map 0 point, y is the output feature map. n Indicated by p 0 As a benchmark, each point in the area covered by the convolution kernel is weighted by summing all the points in the convolution kernel as shown in the above formula, and outputting p 0 The value of y(p 0 ). But the traditional convolution kernel is square and its range of action is also square. However, in complex real-world scenes, especially in low-light conditions, the shape of objects may be distorted or deformed, making it difficult for traditional convolution to capture accurate target features.
[0116] Taking the scene of electricity meter seal detection as an example, in a low-light environment, the image contrast of the electricity meter seal decreases and the details become blurred. Traditional convolution may have difficulty extracting sufficient effective features from such low-quality images. Deformable convolution can more accurately locate and extract seal-related features from limited image information by adaptively adjusting the sampling point positions. Thus, even in low-light conditions, it can capture the weak feature signals of the seal as much as possible and improve the detection accuracy of the seal. The above explains the reason for introducing deformable convolution in the C2f module of the technical solution of this application. The deformable convolution method allows the filter to dynamically adjust its sampling position according to the features of the input image, so as to better adapt to the morphological changes of the target. Specifically, deformable convolution adds a position offset vector Δp to the conventional convolution calculation formula n to adjust the position of the sampling point. The calculation of deformable convolution becomes:
[0117]
[0118] p 0 is the pixel point of the feature map, p 0 + p n is the sampling point of the conventional convolution, p 0 + p n + Δp n is the sampling point of the deformable convolution. However, since the offset Δp n is usually a decimal number, the position p 0 + p n + Δp n after adding the offset is generally a decimal number and does not correspond to the actual pixel points on the input feature map. Therefore, interpolation is needed to obtain the pixel values after offset. Usually, bilinear interpolation is adopted. Finally, the pixel value x(p) after offset is obtained, where the value of p is p 0 + p n + Δp n in the above formula. For bilinear interpolation, its formula is expressed as follows:
[0119]
[0120] where q represents the integer position in the feature map and G represents the bilinear interpolation operation.
[0121] In the DC2f module, compared with the C2f module, all standard convolutions are replaced with deformable convolutions to allow the convolution kernel to dynamically adjust the sampling position according to the content of the input image. After the deformable convolution completes feature extraction, a batch normalization layer (BN) follows, which is used to stabilize and accelerate the training process. Then, the activation function (SiLU) is applied to the normalized output to introduce non-linearity.
[0122] Figure 6 This is a schematic diagram of the structure of DC2f provided by an embodiment of the present application. In Figure 6 the upper part of, the overall structure of DC2f is shown, and DeformConv represents the deformable convolution therein. In Figure 6 the lower part presents the internal structure of the deformable convolution. In DC2f, the input feature map first passes through a deformable convolution layer (with a convolution kernel size of 1×1), and the number of channels of the output feature map doubles. The purpose of this step is to increase the feature expression ability of the model. The convolved features are split into two (i.e., the Split operation is performed), one part enters the Bottleneck module, and the other part directly participates in the subsequent splicing. The split feature maps are processed layer by layer in multiple Bottleneck modules to extract deeper features. It should be noted that the difference between the Bottleneck module in DC2f and the Bottleneck in YOLOv8 is that the convolution is replaced with a deformable convolution. The outputs of all Bottleneck modules and the previously split feature maps are spliced together to increase the diversity of features. Finally, a convolution layer is used to compress the number of channels of the spliced feature map to the required number of output channels to meet the processing requirements of the next step. As Figure 6 shown, in the deformable convolution, offsets represent the offsets.
[0123] Figure 7 This is a schematic diagram of the network structure of another detection model provided by an embodiment of the present application. In Figure 7 the feature extraction network, the feature fusion network, and the prediction output network are shown. As Figure 7 shown, in both the feature extraction network and the feature fusion network, multiple DC2f modules are provided; the DC2f module is a deformable C2f module with a deformable convolution; the positions of the DC2f modules in the network structure of the detection model are the same as the positions of the C2f modules in the network structure of YOLOv8. In addition, in Figure 7 the network structure of the detection model shown, a key improvement introduced above is also shown, that is, a multi-branch fusion module is added. For ease of understanding, the terms belonging to some inherent structures in the YOLOv8 model in the detection model are described below. Conv is convolution, SPPF (Spatial Pyramid Pooling Fast) is a fast version of spatial pyramid pooling, Unsample is an upsampling operation, and Det.Head represents the detection head.
[0124] For the detection model in the present application, model training is required in the early stage. Figure 8A schematic diagram of the training and application of a detection model provided by an embodiment of the present application. The left branch in the figure shows the training of the detection model, and the right branch shows the application of the detection model. In the embodiment of the present application, as mentioned in the relevant introduction of step S102 above, it is necessary to perform image preprocessing on the directly collected images. This is the case in both the training stage and the model application stage because the way of collecting images in the training stage is likely to be the same as that in the model application stage, resulting in the need for preprocessing in both cases.
[0125] In addition, in order to prevent problems such as shadows in low-light environments from affecting the detection performance of the model, the embodiment of the present application hopes to eliminate shadows through image preprocessing and improve the accuracy of small target recognition. Taking the model application stage as an example, the implementation method of preprocessing the initially collected images will be introduced below. It can be understood that in the model training stage, the preprocessing method for the images used for training can refer to the introduction here and will not be elaborated.
[0126] Figure 9 A schematic diagram of the process of preprocessing images in an embodiment of the present application. As Figure 9 shown, the image preprocessing process includes:
[0127] S1021. Convert the initially collected image from the RGB color space to the Lab color space to obtain the first converted image.
[0128] Converting the image from the RGB color space to the Lab color space is mainly because the RGB color space cannot reflect brightness information and is only limited to the color itself, resulting in limited shadow recognition ability. In the Lab color space, the L channel is specifically used to reflect brightness information and can well break through this limitation. In practical applications, converting the RGB value (R, G, B) of each pixel point in the image to the Lab color space requires first converting to the XYZ color space and then to the Lab color space. That is, RGB→XYZ→Lab.
[0129] First, standardize the RGB value to the range of [0, 1] to generate the standardized RGB value, that is, R', G', B'. The formula is as follows:
[0130]
[0131] Further, convert to the XYZ color space, as shown in the following formula:
[0132] X = R′×0.4124 + G′×03576 + B′×0.1805
[0133] Y = R′×0.2126 + G′×0.7152 + B′×0.0722
[0134] Z = R' × 0.0193 + G' × 0.1192 + B' × 0.9505
[0135] Finally, convert to the Lab color space as follows:
[0136] L = 116f(Y / Y n ) - 16
[0137] a = 500(f(X / X n ) - f(Y / Y n ))
[0138] b = 200(f(Y / Y n ) - f(Z / Z n ))
[0139] where X n , Y n , Z n are the coordinates of the white reference point under a given light source, with the values being 0.9505, 1, and 1.0888 respectively. The formula for the function f() is as follows, where t is used to denote the independent variable of the function f():
[0140]
[0141] It can be seen that through the operations of the above formulas, ultimately, the effect of converting each pixel in the image from the RGB color space to the Lab color space can be achieved. When all the pixels in the initially acquired image have completed the above conversion, a new image can be obtained, which is called the first conversion image. In the first conversion image, each pixel carries the L-channel information, a-channel information, and b-channel information of the Lab color space.
[0142] To weaken the adverse effect of the shadow area on the detection accuracy, in the embodiments of the present application, further, the shadow area is identified and determined through the information of the L-channel, a-channel, and b-channel in the converted image, as detailed in step S1022. The L-channel represents the luminance information, and the a-channel and b-channel each carry color information.
[0143] S1022. For the first conversion image, determine the shadow area based on the L-channel, a-channel, and b-channel.
[0144] Specifically, when implementing, the L-channel shadow screening threshold of the first conversion image can be determined first according to the L-channel values of each pixel in the first conversion image, and the b-channel shadow screening threshold of the first conversion image can be determined according to the b-channel values of each pixel in the first conversion image. By determining the shadow screening thresholds based on the luminance channel and the color channel respectively, both the luminance and color information will be considered jointly when identifying the shadow area.
[0145] In the embodiments of the present application, specific conditions for using the above threshold are proposed, namely the first condition. The first condition is that the sum of the average values of the a-channel and the b-channel of each pixel in the first converted image is less than or equal to the maximum value of the color channels of the first converted image. Here, the maximum value of the color channels refers to the maximum value among the a-channel values of each pixel in the image and the b-channel values of each pixel, which can also be understood as the maximum color value of the first converted image.
[0146] If the color information of the first converted image meets the first condition, it is considered that the influence of the shadow in the figure on the a-channel and the b-channel is small. Therefore, only the L-channel is used for shadow determination, that is, the pixels that satisfy L being less than the L-threshold are determined as shadow pixels. Thus, in this case, the shadow area in the first converted image is determined according to the L-channel shadow screening threshold, and the pixel points in the first converted image with the L-channel value of the pixel less than the shadow screening threshold are identified as shadow pixel points, and the connected area of the shadow pixel points is marked as the shadow area.
[0147] If the color information of the first converted image does not meet the first condition, that is, the sum of the average values of the a-channel and the b-channel of each pixel in the first converted image is greater than the maximum value of the color channels of the first converted image, it is considered that the influence of the shadow in the figure on the color channels is large. At this time, not only the L-channel needs to be relied on for shadow determination, but also the b-channel needs to be combined for determination. Specifically, the shadow area in the first converted image is determined according to the L-channel shadow screening threshold and the b-channel shadow screening threshold. If the L-channel value of a certain pixel point in the first converted image is less than the L-channel shadow screening threshold and the b-channel value is less than the b-channel shadow screening threshold, then through the comprehensive judgment of the L-channel and the b-channel, this pixel point is identified as a shadow pixel point, and the connected area of the shadow pixel points is marked as the shadow area.
[0148] It should be noted that after the shadow area is determined, in order to reduce the influence of isolated small pixel points in the image on the detection performance, in the present application, closing operation and opening operation can also be performed after initially determining the shadow area to filter out isolated small pixel points. For example, if the number of connected shadow pixel points is less than 30, then the shadow area constructed by this part of the shadow pixel points can be ignored, and it is very likely to be interference rather than a real shadow.
[0149] Regarding the L-channel shadow screening threshold and the b-channel shadow screening threshold mentioned above, the calculation methods of the two are described below.
[0150] In an alternative implementation, determining the L-channel shadow screening threshold of the first converted image based on the L-channel values of the pixels in the first converted image includes: determining the mean value and the standard deviation of the L-channel of the pixels in the first converted image; multiplying the first proportionality coefficient by the standard deviation of the L-channel to obtain a first product, where the first proportionality coefficient is a positive number less than 1; calculating the difference between the mean value of the L-channel and the first product as the L-channel shadow screening threshold of the first converted image.
[0151] Similarly, in an alternative implementation, determining the b-channel shadow screening threshold of the first converted image based on the b-channel values of the pixels in the first converted image includes: determining the mean value and the standard deviation of the b-channel of the pixels in the first converted image; multiplying the second proportionality coefficient by the standard deviation of the b-channel to obtain a second product, where the second proportionality coefficient is a positive number less than 1; calculating the difference between the mean value of the b-channel and the second product as the b-channel shadow screening threshold of the first converted image.
[0152] As an example, both the first proportionality coefficient and the second proportionality coefficient are one-third. Of course, in other possible implementations, the values of the first proportionality coefficient and the second proportionality coefficient may also be different.
[0153] The standard deviation is a statistic that measures the degree of dispersion of data. In the Lab color space of an image, the standard deviation can reflect the dispersion of the pixel values in that channel relative to the average value. Since L represents luminance information, the standard deviation of the L-channel of the pixels in the shadow area is relatively large. The maximum and minimum values only reflect two extreme cases in the data, and they are easily affected by individual outliers in the image and cannot well represent the pixel value distribution of the entire channel. Therefore, choosing to subtract one-third of the standard deviation from the average value in this application is a measure of trade-off. If the entire standard deviation is subtracted from the average value as the threshold, it may make the threshold too sensitive, resulting in too many pixels being misjudged as shadow or non-shadow areas. By choosing one-third of the standard deviation, while ensuring a certain consideration of the pixel value distribution, the sensitivity of the threshold can be reduced, making the threshold more stable and robust, and better able to adapt to the characteristics of different images.
[0154] S1023. For the shadow area, perform channel value correction in the Lab color space.
[0155] In the embodiments of the present application, in order to weaken the negative impact of shadows on target detection, it is proposed to correct the shadow area with the non-shadow area. Specifically, the average LAB value of the shadow area is obtained, and the average LAB value of the non-shadow area in the first converted image is obtained; the quotient of the average LAB value of the non-shadow area and the average LAB value of the shadow area is calculated as the correction scale factor of the first converted image; the LAB values of the pixels in the shadow area are corrected by using the correction scale factor to obtain the image after the shadow area is corrected. For example, the LAB values of the pixels in the shadow area are multiplied by the above correction scale factor to obtain the corrected LAB values of the pixels in the shadow area. Since the correction scale factor is comprehensively determined based on the average LAB value of the non-shadow area and the average LAB value of the shadow area, the correction scale factor reflects the overall difference between the shadow area and the non-shadow area in terms of the LAB value. Therefore, by using this correction scale factor, the correction of the shadow area can be well achieved, and the similarity between the original shadow area and the non-shadow area can be narrowed.
[0156] S1024. Convert the image after the shadow area is corrected from the Lab color space back to the RGB color space to obtain the second converted image as the input image of the constructed detection model.
[0157] It can be understood that after the shadow area is corrected, the image needs to be converted from the Lab color space to the RGB color space. The intermediate process is to first convert the image from the Lab color space to the XYZ space, and then from the XYZ space to the RGB color space. Since the process of RGB→XYZ→Lab has been described in detail above, as the inverse transformation here, it can be referred to the previous introduction and will not be elaborated.
[0158] In an alternative implementation, converting the image after the shadow area is corrected from the Lab color space back to the RGB color space includes: smoothing the edge of the shadow area in the image after the shadow area is corrected through a median filter to obtain the image after the edge smoothing process; converting the image after the edge smoothing process from the Lab color space back to the RGB color space. The smoothing of the edge improves the quality of the preprocessed image, realizes the shadow elimination scheme based on edge detection, and facilitates the subsequent detection model to more accurately identify and detect small targets.
[0159] In the embodiments of the present application, in order to improve the accuracy and efficiency of object detection, before performing object detection, a series of preprocessing operations are performed on the input image data. In addition to the shadow elimination mentioned above, it may also include size adjustment. Specifically, to ensure that all input images have a unified size to meet the requirements of the network structure and ensure the consistency of model training and inference, the initially acquired images are adjusted to a predetermined resolution. This step is used to resize the images and add padding. First, the current size of the image is obtained, and the scaling ratio is calculated to resize the image while maintaining the aspect ratio. Next, the required padding is calculated, and the padding is evenly distributed on both sides to make the resized image fit the target size. In the present invention, the target size is set to 640x640. The pixel values of the resized image are normalized, that is, the pixel values are mapped from the [0, 255] interval to the [0, 1] interval. This step helps to accelerate model convergence and reduce the impact of image differences under different lighting conditions on the model.
[0160] The above is the introduction to the implementation of image preprocessing. For the training of the model, the following will be introduced in detail. In the present application, in order to train the model, images are also collected for objects of the target type, and image preprocessing is performed to form a training data set. The training data set is divided into a training set, a validation set, and a test set. The training set is used for model training, the validation set is used to verify the generalization of the model and prevent overfitting, and the test set is used to finally evaluate the model performance. In practical applications, the ratio of the training set, the validation set, and the test set can be set to 7:1:2, and the random shuffling method is used to ensure the change of the sample order and increase the randomness of training. The training data set is divided into small batches for training to accelerate the training process. As an example, each small batch contains 32 groups (in other examples, it can also be 64, 128, etc., which can be set according to the video memory of the machine and is not limited here) of preprocessed image data. The model performs forward propagation, loss calculation, backpropagation, and parameter update on each small batch.
[0161] Forward propagation: The model receives a batch of 32 groups (only for example) of images as input and passes them layer by layer through the network layers until a predicted output is generated. During this process, the weights of all layers remain unchanged.
[0162] Loss calculation: The loss value is calculated based on the difference between the predicted output and the true label. The loss function of the detection model consists of a classification loss and a regression loss.
[0163] As an example, the classification loss adopts the Varifocal loss (VFL) to measure the difference between the class label predicted by the model and the true label. The loss function is as follows:
[0164]
[0165] Where \(q\) is the label of the sample, \(p\) is the predicted probability of the model, and \(\alpha\) and \(\gamma\) are hyperparameters. When \(q>0\), positive samples are processed. If the predicted probability \(p\) is close to the actual label \(q\), the loss is small; otherwise, the loss is large. When \(q = 0\), negative samples are processed, and \(p\) γ is down-weighted to reduce the impact of negative samples on the loss. In the scenario of electricity meter seal detection, positive and negative samples can be referred to Figure 10 , in Figure 10 , the left side is a positive sample with a seal on the electricity meter, and the right side is a negative sample without a seal on the electricity meter.
[0166] The regression loss includes CIoU and DFL.
[0167]
[0168] Among them, a learnable hyperparameter \(\alpha\) is introduced, which is used to calculate the diagonal length of the bounding box and the relative position relationship between two bounding boxes (predicted box and ground truth box). The role of \(\alpha\) is to control the convergence speed of CIoU. When \(\alpha = 0\), the calculation result of CIoU is IoU. As \(\alpha\) increases, the calculation result of CIoU gets closer and closer to CIoU. Compared with directly using a fixed value to calculate \(\alpha\), the calculation result is more flexible, enabling the model to have better adaptability to small targets. \(v\) is a parameter used to measure the consistency of aspect ratio and is a difference term reflecting the difference in aspect ratio of the bounding box. The role of \(v\) is to control the width and height of the predicted box to be as close as possible to those of the ground truth box as quickly as possible. \(\rho(b, b gt ) represents the Euclidean distance between the center point of the predicted box \(b\) and the center point of the ground truth box \(b gt . IoU is the standard intersection over union of the predicted box and the ground truth box, and \(c\) represents the diagonal length of the smallest enclosing box of the predicted box and the ground truth box.
[0169] The calculation formulas for the parameters \(\alpha\) and \(v\) in the above formula are as follows:
[0170]
[0171]
[0172] Where \(w\) is the width of the predicted box, \(h\) is the height of the predicted box, \(w gt is the width of the ground truth box, and \(h gt is the height of the ground truth box.
[0173] DFL combines the ideas of Focal Loss and distribution loss, aiming to alleviate the problem of unbalanced positive and negative sample distributions in object detection.
[0174] Loss DFL (S i , S i+1 ) = -((yi+1 -y) log(S i ) + (y - y i ) log(S i+1 ))
[0175] S i and S i+1 are the probability values related to y i and y i+1 respectively. The DFL objective is to increase the probabilities corresponding to the two boundary values of y i and y i+1 . y is the positive sample position, and y i and y i+1 are the two values closest to y. y i <= y <= y i+1 .
[0176] Backpropagation: Once the loss value is determined, the backpropagation algorithm is initiated to propagate the error gradient from the output layer to the input layer along the network. This process calculates the derivative of each layer's weights with respect to the total loss, i.e., the gradient.
[0177] Parameter Update: In this application, the Adam optimizer can be used to update the weights and other parameters in the network based on the calculated gradients. This step aims to minimize the loss function so that the model can better fit the training data in the next iteration.
[0178] During model training, mosaic data augmentation can be disabled in the last 10 epochs to stabilize the training before the end. The initial learning rate is set to 0.01, and the final learning rate is set to 0.0001 to adjust the change of the learning rate over time. The parameter beta1 of the Adam optimizer is set to 0.937, and the L2 regularization term penalty is set to 0.0005. The number of epochs for learning rate warm-up is set to 3, gradually increasing the learning rate from a low value to the initial learning rate. The initial momentum in the warm-up stage is set to 0.8, and the bias parameter is set to 0.1, which helps to stabilize the model training in the initial epochs. The weight of the CIoU loss component in the loss function is set to 7.5, the VFL loss is set to 0.5, and the DFL loss is set to 1.5.
[0179] Taking the detection of the electricity meter seal as an example, the effects of seal detection are compared through several groups of images. In the shown effect comparison, meter represents the electricity meter, monitor represents the display screen of the electricity meter, number represents the barcode identification of the electricity meter, and lock represents the electricity meter seal. The confidence level of the decimal classification beside the classification label. Among them, on the left is YOLO vThe operation result on the right is the operation result of the detection model proposed in this application. It should be noted that in several subsequent comparison diagrams of effects, words such as "meter", "monitor", "number", "lock", etc. may not be fully displayed, and the corresponding confidence levels may also not be fully displayed. The reason for these situations is that the display areas of the detection results overlap or intersect. When observing the comparison effects, only the words and values that are fully displayed and not blocked can be observed.
[0180] In Figure 11 A comparison of the effects of the first group of lead seal detections is shown. The model with the YOLOv8 structure only detected the lead seals on 5 electric meters, while the detection model proposed in this application detected the lead seals on 7 electric meters. And according to the displayed confidence levels, it can be seen that the confidence levels of the detected lead seals have increased. In Figure 12 In the comparison of the effects of the second group of lead seal detections shown, the detection model proposed in this application successfully detected a very small target, that is, the lead seal of the electric meter. While the model with the YOLOv8 structure could not detect it. In Figure 13 , Figure 14 , Figure 15 In the comparisons of the effects of the third, fourth, and fifth groups of lead seal detections shown, similarly, it is shown that the detection effect of the detection model in this application is more complete and the confidence level is higher. In Figure 16 In the comparison of the effects of the sixth group of lead seal detections shown, it can be seen that the model with the YOLOv8 structure misidentified a small section of wire as a lead seal, while the detection model in this application successfully identified that the small section of wire was not a lead seal. Combining the Figures 11 to 16 shown comparison effects, it is not difficult to see that the embodiment of this application applies an improved detection model, which has the performance and effects of missing detection and misdetection of small targets in low-light environments.
[0181] Based on the object detection method introduced in the previous embodiments, this application correspondingly also provides an object detection device. The implementation of this device will be described below with reference to the accompanying drawings. As Figure 17 , it is a schematic structural diagram of an object detection device provided by an embodiment of this application. As Figure 17 shown, the object detection device includes:
[0182] An image acquisition module 171, configured to acquire an initial captured image of an object of a target type in a detection scene; the first additional object of the object is the detection target, and the size of the first additional object is smaller than the size of the main structure of the object;
[0183] The image preprocessing module 172 is used to perform image preprocessing on the initial acquired image to construct the input image of the detection model; the network structure of the detection model is a network structure improved based on YOLOv8, which includes a feature extraction network, a feature fusion network, and a prediction output network; a multi-branch fusion module is set in the feature fusion network, and the multi-branch fusion module includes multiple parallel branches, and the multiple parallel branches are respectively used to capture information of different scales in the input feature map of the multi-branch fusion module;
[0184] The target detection module 173 is used to input the input image into the detection model, and perform image analysis through the detection model to obtain the image target detection result output by the detection model; the image target detection result is used to indicate whether the object has or does not have the first additional object.
[0185] The multi-branch fusion module has multiple parallel branches, which are respectively used to capture information of different scales in the input feature map of the multi-scale fusion module. Since the multi-branch fusion module in the detection model can provide multi-scale receptive fields, the application of the multi-branch fusion module extracts image features from multiple scales, which can cover various scale changes of small targets in the acquired image. In this way, the ability of the detection model to extract and recognize small target features is greatly enhanced, and the detection performance of the model for small targets in low-light environments is improved. The technical solution of this application also has a high application prospect in low-light environments. Through experimental verification, this solution can reduce the risk of missed detection and false detection of small targets in low-light environments and achieve target detection more accurately.
[0186] In an optional implementation manner, during the process of the target detection module 173 performing image analysis through the detection model, the multi-branch fusion module in the detection model is used for:
[0187] Perform a first convolution operation on the input feature map to obtain a first convolution result;
[0188] Divide the information in the first convolution result into first part information and second part information according to channels;
[0189] Through the multiple parallel branches, respectively perform feature extraction of different scales on the first part information to obtain multiple feature extraction results; the multiple feature extraction results correspond to the multiple parallel branches one by one;
[0190] Perform channel information superposition on the multiple feature extraction results and the second part information to obtain a superposition result of information;
[0191] Perform a second convolution operation on the information superposition result to obtain a second convolution result as the output feature map of the multi-branch fusion module.
[0192] In an alternative implementation, the multi-branch fusion module is specifically configured to:
[0193] Determine the total number of channels of the first convolution result, and determine the multi-scale perception proportionality coefficient of the multi-branch fusion module; the multi-scale perception proportionality coefficient is a positive number less than 1;
[0194] Based on the total number of channels, extract information of corresponding proportions of channels from the first convolution result as the first part of information based on the multi-scale perception proportionality coefficient, and use the remaining information in the first convolution result as the second part of information.
[0195] In an alternative implementation, the multiple parallel branches include a first branch, a second branch, and a third branch, where the receptive field of the first branch is smaller than the receptive field of the second branch, and the receptive field of the third branch is larger than the receptive field of the second branch;
[0196] The first branch includes a standard convolutional layer with a convolutional kernel of a first size;
[0197] The second branch includes a first convolutional layer, a second convolutional layer, and a third convolutional layer, where the first convolutional layer is used to convolve the horizontal features of the image, and the third convolutional layer is used to convolve the vertical features of the image; the second convolutional layer is located between the first convolutional layer and the third convolutional layer, and the second convolutional layer is a standard convolutional layer with a convolutional kernel of a second size, and the second size is larger than the first size; the third branch includes a module based on a spatial attention mechanism and / or a channel attention mechanism.
[0198] In an alternative implementation, the third branch specifically includes a dual-domain channel attention module and a frequency-based spatial attention module; the dual-domain channel attention module is located at the front end of the frequency-based spatial attention module, and the output of the dual-domain channel attention module is used as the input of the frequency-based spatial attention module.
[0199] In an alternative implementation, multiple DC2f modules are provided in both the feature extraction network and the feature fusion network; the DC2f module is a deformable C2f module with deformable convolution; the positions of the DC2f modules in the network structure of the detection model are the same as the positions of the C2f modules in the network structure of YOLOv8.
[0200] In an alternative implementation, the image preprocessing module 172 includes:
[0201] A first color space conversion unit for converting the initial captured image from the RGB color space to the Lab color space to obtain a first converted image;
[0202] A shadow region determination unit, configured to determine a shadow region for the first converted image based on the L channel, the a channel, and the b channel;
[0203] A correction unit, configured to perform channel value correction on the shadow region in the Lab color space;
[0204] A second color space conversion unit, configured to convert the image after the shadow region is corrected from the Lab color space back to the RGB color space, and obtain a second converted image as the input image for constructing the detection model.
[0205] In an optional implementation manner, the shadow region determination unit is specifically configured to:
[0206] Determine an L channel shadow screening threshold of the first converted image according to the L channel values of the pixels in the first converted image, and determine a b channel shadow screening threshold of the first converted image according to the b channel values of the pixels in the first converted image;
[0207] If the color information of the first converted image satisfies a first condition, determine the shadow region in the first converted image according to the L channel shadow screening threshold; if the color information of the first converted image does not satisfy the first condition, determine the shadow region in the first converted image according to the L channel shadow screening threshold and the b channel shadow screening threshold; the first condition is that the sum of the a channel mean value and the b channel mean value of each pixel in the first converted image is less than or equal to the maximum value of the color channels of the first converted image; the maximum value of the color channels is the maximum value among the a channel values of each pixel and the b channel values of each pixel.
[0208] In an optional implementation manner, the shadow region determination unit is specifically configured to:
[0209] Determine the L channel mean value and the L channel standard deviation of each pixel in the first converted image;
[0210] Multiply a first proportionality coefficient by the L channel standard deviation to obtain a first product; the first proportionality coefficient is a positive number less than 1;
[0211] Calculate the difference between the L channel mean value and the first product as the L channel shadow screening threshold of the first converted image;
[0212] The determining the b channel shadow screening threshold of the first converted image according to the b channel values of the pixels in the first converted image includes:
[0213] Determine the b channel mean value and the b channel standard deviation of each pixel in the first converted image;
[0214] Multiply the second proportionality coefficient by the standard deviation of the b channel to obtain a second product; the second proportionality coefficient is a positive number less than 1;
[0215] Calculate the difference between the mean value of the b channel and the second product as the b channel shadow screening threshold of the first converted image.
[0216] In an alternative implementation, the correction unit is specifically configured to:
[0217] Obtain the average LAB value of the shadow area and obtain the average LAB value of the non-shadow area in the first converted image;
[0218] Calculate the quotient of the average LAB value of the non-shadow area and the average LAB value of the shadow area as the correction scale factor of the first converted image;
[0219] Use the correction scale factor to correct the LAB values of the pixels in the shadow area to obtain an image with the shadow area corrected.
[0220] In an alternative implementation, the second color space conversion unit is specifically configured to:
[0221] Smooth the edges of the shadow area in the image with the shadow area corrected by a median filter to obtain an image with the edges smoothed;
[0222] Convert the image with the edges smoothed from the Lab color space back to the RGB color space.
[0223] Based on the apparatus for object detection and the method for object detection introduced in the foregoing embodiments, an apparatus for object detection is further provided in an embodiment of the present application. The apparatus includes: a memory and a processor; the memory is used to store a computer program; the processor is used to run the computer program, and when the computer program runs, it executes some or all of the steps in the method for object detection introduced in the method embodiment.
[0224] It should be noted that the embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content. The device and equipment embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0225] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for target detection, characterized in that: include: Obtaining an initial acquisition image of an object of the target type in the detection scene; The first additional object of the object is a detection target, and the size of the first additional object is smaller than the size of the main structure of the object; Performing image preprocessing on the initial collected image to construct an input image of a detection model; The network structure of the detection model is an improved network structure based on YOLOv8, which includes a feature extraction network, a feature fusion network and a prediction output network; A multi-branch fusion module is provided in the feature fusion network, and the multi-branch fusion module includes a plurality of parallel branches, and the plurality of parallel branches are respectively used to capture information of different scales in the input feature map of the multi-branch fusion module; After the input image is input into the detection model, the detection model is used to perform image analysis to obtain an image target detection result output by the detection model; The image object detection result is used to indicate whether the object has the first additional object or not.
2. The method according to claim 1, characterized in that In the process of performing image analysis by the detection model, the processing steps of the multi-branch fusion module include: Performing a first convolution operation on the input feature map to obtain a first convolution result; Dividing the information in the first convolution result into a first part of information and a second part of information according to channels; Through the multiple parallel branches, feature extraction of different scales is performed on the first part of information respectively to obtain multiple feature extraction results; the multiple feature extraction results correspond to the multiple parallel branches one by one; Perform channel information superposition on the multiple feature extraction results and the second part of information to obtain an information superposition result; A second convolution operation is performed on the information superposition result to obtain a second convolution result as an output feature map of the multi-branch fusion module.
3. The method according to claim 2, characterized in that The dividing the information in the first convolution result into a first part of information and a second part of information according to the channel includes: Determine the total number of channels of the first convolution result, and determine the multi-scale perception scale coefficient of the multi-branch fusion module; the multi-scale perception scale coefficient is a positive number less than 1; Based on the total number of channels, information of channels of a corresponding proportion is extracted from the first convolution result based on the multi-scale perception scale coefficient as the first part of information, and the remaining information in the first convolution result is used as the second part of information.
4. The method according to any one of claims 1 to 3, characterized in that The multiple parallel branches include a first branch, a second branch and a third branch, wherein the receptive field of the first branch is smaller than the receptive field of the second branch, and the receptive field of the third branch is larger than the receptive field of the second branch; The first branch includes a standard convolutional layer having a convolutional kernel of a first size; The second branch includes a first convolution layer, a second convolution layer and a third convolution layer, wherein the first convolution layer is used to convolve the horizontal features of the image, and the third convolution layer is used to convolve the vertical features of the image; the second convolution layer is located between the first convolution layer and the third convolution layer, and the second convolution layer is a standard convolution layer with a convolution kernel of a second size, and the second size is larger than the first size; the third branch includes a module based on a spatial attention mechanism and / or a channel attention mechanism.
5. The method according to claim 4, characterized in that The third branch specifically includes a dual-domain channel attention module and a frequency-based spatial attention module; the dual-domain channel attention module is located at the front end of the frequency-based spatial attention module, and the output of the dual-domain channel attention module serves as the input of the frequency-based spatial attention module.
6. The method according to any one of claims 1 to 3, characterized in that A plurality of DC2f modules are provided in the feature extraction network and the feature fusion network; the DC2f module is a deformable C2f module having a deformable convolution; and the position of each DC2f module in the network structure of the detection model is the same as the position of the C2f module in the network structure of YOLOv8.
7. The method according to any one of claims 1 to 3, characterized in that The performing image preprocessing on the initial collected image to construct an input image of the detection model includes: Convert the initial acquired image from the RGB color space to the Lab color space to obtain a first converted image; For the first converted image, determining a shadow area based on the L channel, the a channel, and the b channel; For the shadow area, performing channel value correction in Lab color space; The image after the shadow area correction is converted from the Lab color space back to the RGB color space to obtain a second converted image as the input image of the constructed detection model.
8. The method according to claim 7, characterized in that The step of determining the shadow area based on the L channel, the a channel and the b channel for the first converted image comprises: Determine an L channel shadow screening threshold of the first conversion image according to an L channel value of each pixel in the first conversion image, and determine a b channel shadow screening threshold of the first conversion image according to a b channel value of each pixel in the first conversion image; If the color information of the first conversion image satisfies a first condition, a shadow area in the first conversion image is determined according to the L channel shadow screening threshold; if the color information of the first conversion image does not satisfy the first condition, a shadow area in the first conversion image is determined according to the L channel shadow screening threshold and the b channel shadow screening threshold; the first condition is: the sum of the a channel mean value and the b channel mean value of each pixel in the first conversion image is less than or equal to the color channel maximum value of the first conversion image; the color channel maximum value is the maximum value of the a channel value of each pixel and the b channel value of each pixel.
9. The method according to claim 8, characterized in that The determining the L channel shadow screening threshold of the first converted image according to the L channel value of each pixel in the first converted image comprises: Determine an L channel mean and an L channel standard deviation of each pixel in the first converted image; Multiplying the first proportionality coefficient by the L channel standard deviation to obtain a first product; the first proportionality coefficient is a positive number less than 1; Calculating a difference between the L channel mean and the first product as an L channel shadow screening threshold of the first converted image; The determining the b-channel shadow screening threshold of the first converted image according to the b-channel value of each pixel in the first converted image comprises: Determining a b-channel mean and a b-channel standard deviation of each pixel in the first converted image; Multiplying the second proportional coefficient by the b channel standard deviation to obtain a second product; the second proportional coefficient is a positive number less than 1; A difference between the b channel mean and the second product is calculated as a b channel shadow screening threshold of the first converted image.
10. The method according to any one of claims 7 to 9, characterized in that: The step of performing channel value correction on the shadow area in the Lab color space includes: Obtaining an average LAB value of the shadow area, and obtaining an average LAB value of a non-shadow area in the first converted image; Calculate the quotient of the average LAB value of the non-shadow area and the average LAB value of the shadow area as a correction scale factor of the first converted image; The LAB value of each pixel in the shadow area is corrected using the correction scale factor to obtain an image after the shadow area is corrected.
11. The method according to any one of claims 7 to 9, characterized in that: The step of converting the image corrected in the shadow area from the Lab color space back to the RGB color space comprises: Smoothing the edges of the shadow area in the image after the shadow area correction by using a median filter to obtain an image after edge smoothing; The edge-smoothed image is converted from the Lab color space back to the RGB color space.
12. A target detection device, characterized in that: include: An image acquisition module is used to acquire an initial acquisition image of an object of a target type in a detection scene; The first additional object of the object is a detection target, and the size of the first additional object is smaller than the size of the main structure of the object; An image preprocessing module is used to perform image preprocessing on the initial collected image to construct an input image of the detection model; the network structure of the detection model is an improved network structure based on YOLOv8, which includes a feature extraction network, a feature fusion network and a prediction output network; A multi-branch fusion module is provided in the feature fusion network, and the multi-branch fusion module includes a plurality of parallel branches, and the plurality of parallel branches are respectively used to capture information of different scales in the input feature map of the multi-branch fusion module; The target detection module is used to perform image analysis through the detection model after the input image is input into the detection model, so as to obtain the image target detection result output by the detection model; The image object detection result is used to indicate whether the object has the first additional object or not.
13. A target detection device, characterized in that: include: Memory and processor; The memory is used to store computer programs; The processor is used to run the computer program, and the computer program executes the steps of the method according to any one of claims 1 to 11 when running.