Worker safety helmet detection method for power construction scene

By constructing and labeling the image data set of the power construction scenario, and training the object detection model containing multi-scale channel interaction module, attention mechanism and Ghost convolution, the problems of low efficiency and high cost of safety helmet detection in the power construction scenario are solved, and high accuracy safety helmet detection is achieved, reducing safety risks.

CN119942584APending Publication Date: 2025-05-06国网甘肃省电力公司金昌供电公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411762959.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, safety helmet detection in power construction scenarios has problems such as low detection efficiency, high cost and insensitive to complex backgrounds, and traditional object detection algorithms are not sensitive to small object detection.

Method used

A image data set for power construction scenarios is constructed, and annotated and enhanced. The object detection model is trained using the preprocessed data set. The model includes a multi-scale channel interaction module, introducing attention mechanisms and Ghost convolutions, and is improved based on the YOLOv8 object detection model.

Benefits of technology

It improves the accuracy of safety helmet inspection in power construction scenarios, saves supervision and inspection costs, reduces safety risks at construction sites, and promotes the development of the power industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942584A_ABST
    Figure CN119942584A_ABST
Patent Text Reader

Abstract

The invention provides a worker safety helmet detection method for an electric power construction scene. The method comprises the following steps: constructing an image data set for the electric power construction scene, and labeling and enhancing the image data set to obtain a preprocessed data set; and training the constructed target detection model by using the preprocessed data set, and detecting whether the image of the current power construction site contains the worker helmet or not after training is completed to obtain a detection result. The target detection model is based on a YOLOv8 target detection model, the target detection model achieves the purposes of enhancing the model performance and improving the detection accuracy by introducing a multi-scale channel information interaction module, an attention mechanism and Ghost convolution, and is beneficial to improving the safety helmet detection accuracy in a power construction scene; the personal safety of power construction personnel can be further guaranteed, the supervision and detection cost is saved, and the safety risk of a construction site can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of engineering safety, and in particular relates to a method for detecting workers' safety helmets in electric power construction scenarios. Background Art

[0002] As an efficient protective tool, hard hats can minimize head injuries in the event of dangerous accidents, so as to protect the safety of power grid personnel. At present, most of the actual hard hat detection relies on human eye supervision, which has problems such as low detection efficiency and high cost. The monitoring method with manual supervision is out of date. The target detection technology combined with computer vision is gradually developing. However, the traditional target detection algorithm has problems such as insensitivity to complex backgrounds and insensitivity to small target detection of hard hats. A fast and reliable method for detecting hard hats for workers in power construction scenarios is needed. Summary of the invention

[0003] In order to solve the above problems existing in the prior art, the present invention provides a worker helmet detection method for power construction scenes. The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0004] A method for detecting worker helmets in power construction scenarios includes:

[0005] S100, constructing an image dataset for power construction scenes, and annotating and enhancing the image dataset to obtain a preprocessed dataset;

[0006] S200, using the preprocessed data set to train the constructed target detection model to obtain a trained target detection model; wherein the target detection model includes a multi-scale channel interaction module, and the multi-scale channel interaction module introduces an attention mechanism and a Ghost convolution;

[0007] S300, using the trained target detection model to detect whether the image of the current power construction site contains workers' safety helmets, and obtain a detection result.

[0008] Beneficial effects:

[0009] The present invention provides a method for detecting worker helmets for electric power construction scenes, constructs an image data set for electric power construction scenes, and annotates and enhances the image data set to obtain a preprocessing data set; uses the preprocessing data set to train the constructed target detection model to obtain a trained target detection model; wherein the target detection model includes a multi-scale channel interaction module, and the multi-scale channel interaction module introduces an attention mechanism and a Ghost convolution; uses the trained target detection model to detect whether the image of the current electric power construction site contains a worker helmet, and obtains a detection result. The present invention is based on the YOLOv8 target detection model, and by introducing a multi-scale channel information interaction module, an attention mechanism and a Ghost convolution, the purpose of enhancing model performance and improving detection accuracy is achieved, which is conducive to improving the accuracy of helmet detection in electric power construction scenes, can further ensure the personal safety of electric power construction personnel, and saves supervision and detection costs, which is conducive to reducing safety risks at construction sites, thereby promoting the development of the electric power industry, and has important application value.

[0010] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 It is a flow chart of a method for detecting worker safety helmets in electric power construction scenarios provided by the present invention;

[0012] Figure 2 is a schematic diagram of a multi-scale channel interaction module provided by the present invention;

[0013] Figure 3 It is a schematic diagram of the detection results provided by the present invention. DETAILED DESCRIPTION

[0014] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.

[0015] like Figure 1 As shown, the present invention provides a method for detecting worker helmets in electric power construction scenarios, comprising:

[0016] S100, constructing an image dataset for power construction scenes, and annotating and enhancing the image dataset to obtain a preprocessed dataset;

[0017] S200, using the preprocessed data set to train the constructed target detection model to obtain a trained target detection model; wherein the target detection model includes a multi-scale channel interaction module (Multi-scale Channel Information, MCI), and the multi-scale channel interaction module introduces an attention mechanism and a Ghost convolution;

[0018] Wherein, the target detection model adopts the YOLOv8 target detection model, and the YOLOv8 target detection model also includes a backbone network (backbone), a neck network (neck) and an output end (head). The multi-scale channel interaction module is arranged between the backbone network and the neck network; the backbone network uses the Ghost convolution method to convolve the input image to obtain a feature map; the neck network is used to fuse the features output by the multi-scale channel interaction module, and the neck network includes a feature pyramid network and a path aggregation network; the output end detects the result and outputs the detection result; the backbone network includes a CBS module, a CSP module and an SPPF (spatial pyramid pooling module) that introduce Ghost convolution, and the CBS module is a combination of a convolution module, a batch normalization module and a Sigmoid activation function. The neck network is mainly used for feature fusion, which includes a feature pyramid network (FPN) and a path aggregation network (PAN). The output end (head) is responsible for detecting the result and outputting the detection result. The multi-scale channel information interaction module realizes channel dependence across different scales by establishing information interaction between different channels.

[0019] S300, using the trained target detection model to detect whether the image of the current power construction site contains workers' safety helmets, and obtain a detection result.

[0020] In a specific embodiment of the present invention, S100 includes:

[0021] S110, downloading a public electric power construction site image dataset from the Internet, and combining it with images collected at the electric power construction site to form an image dataset for electric power construction scenes; wherein the image dataset includes images of on-site personnel wearing and not wearing safety helmets;

[0022] S120, using LabelImg software to classify and label the images in the image dataset, so as to label the helmet target as hat with a label box, and label the on-site personnel as person with a label box, thereby obtaining a labeled image dataset;

[0023] S130, performing image enhancement on the images in the image dataset carrying the marker by means of rotation, cropping, exposure and denoising to obtain a preprocessed dataset.

[0024] Image augmentation of images in labeled image datasets using rotation, cropping, exposure, and denoising can increase the number of samples and improve the generalization ability of subsequent models.

[0025] In the specific implementation, the constructed helmet dataset has a total of 9321 images. The target boxes in the dataset are annotated into two categories: "hat" and "person".

[0026] In a specific embodiment of the present invention, reference Figure 2 , the multi-scale channel interaction module includes: a first CBS module, a second CBS module, a third CBS module, a multi-scale convolution channel, a first splicing module, an SE attention mechanism module, a second splicing module and a fourth CBS module; the multi-scale interaction module includes four channels, each channel corresponds to a different convolution kernel;

[0027] Among them, the input of the first CBS module is used to input the image and is connected to the input of the third CBS module. The output of the first CBS module is linked to the input of the second CBS module. The output of the second CBS module is connected to the multi-scale convolution channel. The output of the multi-scale convolution channel is connected to the input of the first splicing module. The output of the first splicing module is connected to the input of the SE attention mechanism module. The output of the SE attention mechanism module is multiplied by the output of the first splicing module after the softmax function, and then multiplied with the output of the first CBS module, and then sent to the second splicing module together with the output of the third CBS module. The output of the second splicing module is connected to the input of the fourth CBS module, and the output of the fourth CBS module is used to output the feature map.

[0028] In a specific implementation of the present invention, S200 includes:

[0029] S210, dividing the image dataset into a training set, a validation set, and a test set in a ratio of 8:1:1;

[0030] S220, selecting an input image from the training set and inputting it into the constructed object detection model, so as to output a prediction result using the object detection model, wherein the prediction result includes a prediction box marking a helmet in the input image;

[0031] S230, calculating a loss function using a prediction box and an annotation box in the input image in the prediction result;

[0032] S240, using the loss function to determine whether the training end condition is met, if not, adjusting the weight of the target detection model, and returning to S220 until the training end condition is met, to obtain a trained target detection model.

[0033] After the training is completed, the present invention obtains the model weight with the best performance and saves it as a file best.pt for subsequent use.

[0034] In a specific implementation of the present invention, S220 includes:

[0035] S221, selecting an input image from the training set and inputting it into the backbone network, so as to extract features of the input image using the backbone network, and obtaining an input feature map X with a size of c×h×w;

[0036] S222, the input feature map X of size c×h×w is processed by the first CBS module to obtain a feature map of size Feature map X1;

[0037] S223, the size is The feature map X1 is processed by the second CBS module to produce a size of The feature map F;

[0038] S224, the The feature map F is divided into four parts, which are convolved through four scale convolution channels to obtain a size of Four feature maps F0~F3;

[0039] S225, the size is The four feature maps F0~F3 of are processed by the first splicing module to generate a multi-scale channel feature with a size of Feature map K1;

[0040] S225, weighting the feature map K1 through the CA attention mechanism module to generate a refined feature map K2 with attention weights.

[0041] The present invention introduces a CA attention mechanism module into the multi-scale channel information interaction module to improve the model's detection accuracy for small targets. The attention mechanism is a way of processing visual information from the human brain. While paying attention to important features, it can not only suppress and remove other unnecessary features, but also convert the spatial information in the image accordingly to extract key information. The CA self-attention mechanism encodes channel relationships and long-term dependencies through precise location information, and is composed of Coordinate information embedding and Coordinate attention generation. The refined feature map K2 weighted by the CA attention mechanism module is expressed as:

[0042]

[0043]

[0044] Where: b(·)——global average pooling operation; a1(·), a2(·)——two fully connected layers; σ(·)——Sigmoid activation function; δ(·)——ReLU activation function.

[0045] The specific formulas of Sigmoid activation function and ReLU activation function are as follows:

[0046]

[0047]

[0048] Then the obtained weighted feature map K2 is added to the feature map X1 to obtain the feature map X2. X2 and X1 are then concatenated to generate the output feature map Y of the MCI module.

[0049] S226, adding the refined feature map K2 to the feature map X1 to obtain the feature map X2, and outputting the feature map X2 and the feature map X1 to a feature map Y with a size of c×h×w through a fourth CBS module;

[0050] S227, passing the feature map Y through the neck network and the output end to obtain a prediction result of the input image.

[0051] In a specific embodiment of the present invention, S225 includes:

[0052] S2251, the CA attention mechanism module encodes the feature map K1 in the height and width directions respectively to generate a feature map in the height direction and a feature map in the width direction; the feature map in the height direction and the feature map in the width direction are respectively expressed as:

[0053]

[0054]

[0055]

[0056]

[0057] Where: z h Represents the feature map in the height direction; z w Represents the feature map in the width direction; Indicates the encoding value output by the cth channel in the height direction; Indicates the encoding value output by the cth channel in the width direction; w, h - the width and height of the image; x c (h, i) represents the value of feature map K1 at point (h, i) in the cth channel;

[0058] S2252, concatenate the feature maps in the height direction and the width direction, perform 1×1 convolution, and activate them using the Sigmoid function to obtain an intermediate feature map f of size c×(h+w) containing width encoding and height encoding information; the intermediate feature map f is expressed as:

[0059] f = δ(conv 1×1 ([z h , z w ]))

[0060] The specific formula of the Sigmoid activation function is as follows:

[0061]

[0062] S2253, split the intermediate feature map f to generate two independent tensors f h ∈R C×H and f w ∈R C ×W ;

[0063] S2254, unify two independent tensors f using two 1×1 convolution operations respectively h ∈R C×H and f w ∈R C×W The number of channels is combined with the sigmoid activation function to obtain the attention vector; the attention vector is expressed as:

[0064]

[0065]

[0066] S2255, using the attention vector to perform attention weighting on the feature map K1 to obtain the refined features on each channel; the refined features on each channel are expressed as:

[0067]

[0068] in, is the refined feature on the cth channel, is the feature map K1 on the cth channel, is the height-direction attention vector on the c-th channel, The width-wise attention vector on the c-th channel.

[0069] S2256, the refined features on each channel are combined into a refined feature map K2.

[0070] In a specific implementation of the present invention, the CBS module introducing Ghost convolution uses the Ghost convolution method to convolve the input image, and the process includes:

[0071] Extract features from the input image to obtain multiple features;

[0072] Perform conventional convolution on a part of the multiple features to obtain the intrinsic feature map;

[0073] Use a 3×3 convolution kernel to perform Ghost convolution on the remaining parts of multiple features to obtain a Ghost feature map;

[0074] The intrinsic feature map and the Ghost feature map are spliced ​​to obtain a CBS module spliced ​​feature map.

[0075] In the backbone part of YOLOv8, the Ghost convolution method is used to replace the original convolution method of the model. Ghost convolution is divided into three steps: regular convolution, generating Ghost feature map and splicing feature maps. First, part of the input feature map is convolved with regular convolution to obtain the intrinsic feature map, thereby reducing most of the calculation amount. Then, the remaining feature map is convolved with a 3×3 convolution kernel, and the obtained feature map becomes the Ghost feature map. Finally, the intrinsic feature map obtained in the first step is spliced ​​with the Ghost feature map obtained in the second step to obtain the convolved feature map. The use of the Ghost convolution method can significantly reduce the number of parameters and calculations, save computer resources and improve the detection speed of the model.

[0076] In a specific embodiment of the present invention, the loss function includes the confidence regression loss L obj , category loss L cls and the localization loss L loc ;

[0077] The confidence regression loss is calculated using binary cross entropy and is expressed as:

[0078]

[0079] In the formula, λ obj , noobj Correspondingly, the coefficients of confidence regression error with and without targets; S- represents the number of grids used to divide the image in the backbone network; B represents the number of prediction boxes generated for each grid in the feature map; Indicates whether there is an object in the prediction box of the ith network. When the intersection and union ratio of the annotation box and the prediction box is greater than the set threshold, set is 1, otherwise it is 0; C i Indicates the confidence of the prediction box; Indicates the confidence of the annotation box.

[0080] The positioning loss is calculated using CIOU, which is expressed as follows:

[0081]

[0082]

[0083]

[0084]

[0085] Where bgt represents the coordinates of the center point of the annotation box; b represents the coordinates of the center point of the prediction box; c represents the diagonal length of the minimum circumscribed rectangle of the annotation box and the prediction box; ρ represents the Euclidean distance between the center points of the annotation box and the prediction box; α represents the weight parameter; v represents the aspect ratio of the prediction box, and w gt Indicates the height of the annotation box, h gt represents the height of the annotation box, w represents the width of the prediction box, h represents the height of the prediction box; αv represents the penalty item set, the purpose of which is to make the width and height of the prediction box approach the annotation box faster.

[0086] The effect of the present invention is explained below through simulation experiments.

[0087] The hyperparameters of the training process are set as follows: the number of epochs is 500, the optimizer selects the SGD optimizer, the initial learning rate is set to 0.01, the batch size is set to 16, and the experimental environment configuration used is shown in the following table.

[0088] Table 1 Experimental environment configuration table

[0089]

[0090] Use the best weight best.pt obtained from training to detect helmets on the test set. The detection accuracy is 95.6% and the recall is 94.1%. It has a good detection effect. Some of the test results are shown in the attached figure. Figure 3 shown.

[0091] The implementation of the worker safety helmet detection method for electric power construction scenes in the present invention is beneficial to improving the safety helmet detection accuracy in electric power construction scenes, can further ensure the personal safety of electric power construction personnel, and saves supervision and detection costs, which is beneficial to reducing safety risks at construction sites, thereby promoting the development of the electric power industry and has important application value.

[0092] It is worth noting that the terms "first" and "second" in the present invention are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0093] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps.

[0094] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.

Claims

1. A worker helmet detection method for electric power construction scenes, characterized in that: include: S100, constructing an image dataset for power construction scenes, and annotating and enhancing the image dataset to obtain a preprocessed dataset; S200, using the preprocessed data set to train the constructed target detection model to obtain a trained target detection model; wherein the target detection model includes a multi-scale channel interaction module, and the multi-scale channel interaction module introduces an attention mechanism and a Ghost convolution; S300, using the trained target detection model to detect whether the image of the current power construction site contains workers' safety helmets, and obtain a detection result.

2. The method for detecting worker helmets in electric power construction scenarios according to claim 1 is characterized in that: S100 includes: S110, downloading a public electric power construction site image dataset from the Internet, and combining it with images collected at the electric power construction site to form an image dataset for electric power construction scenes; wherein the image dataset includes images of on-site personnel wearing and not wearing safety helmets; S120, using LabelImg software to classify and label the images in the image dataset, so as to label the helmet target as hat with a label box, and label the on-site personnel as person with a label box, thereby obtaining a labeled image dataset; S130, performing image enhancement on the images in the image dataset carrying the marker by means of rotation, cropping, exposure and denoising to obtain a preprocessed dataset.

3. The method for detecting worker helmets in electric power construction scenes according to claim 1 is characterized in that: The target detection model adopts the YOLOv8 target detection model, which also includes a backbone network, a neck network and an output end. The multi-scale channel interaction module is arranged between the backbone network and the neck network; the backbone network uses the Ghost convolution method to convolve the input image to obtain a feature map; the neck network is used to fuse the features output by the multi-scale channel interaction module, and the neck network includes a feature pyramid network and a path aggregation network; the output end detects the result and outputs the detection result.

4. The method for detecting worker helmets in electric power construction scenes according to claim 3 is characterized in that: The multi-scale channel interaction module includes: a first CBS module, a second CBS module, a third CBS module, a multi-scale convolution channel, a first splicing module, an SE attention mechanism module, a second splicing module and a fourth CBS module; the multi-scale interaction module includes four channels, each channel corresponds to a different convolution kernel; Among them, the input of the first CBS module is used to input the image and is connected to the input of the third CBS module. The output of the first CBS module is linked to the input of the second CBS module. The output of the second CBS module is connected to the multi-scale convolution channel. The output of the multi-scale convolution channel is connected to the input of the first splicing module. The output of the first splicing module is connected to the input of the SE attention mechanism module. The output of the SE attention mechanism module is multiplied by the output of the first splicing module after the softmax function, and then multiplied with the output of the first CBS module, and then sent to the second splicing module together with the output of the third CBS module. The output of the second splicing module is connected to the input of the fourth CBS module, and the output of the fourth CBS module is used to output the feature map.

5. The method for detecting worker helmets in electric power construction scenarios according to claim 4 is characterized in that: S200 includes: S210, dividing the image dataset into a training set, a validation set, and a test set in a ratio of 8:1:1; S220, selecting an input image from the training set and inputting it into the constructed object detection model, so as to output a prediction result using the object detection model, wherein the prediction result includes a prediction box marking a helmet in the input image; S230, calculating a loss function using a prediction box and an annotation box in the input image in the prediction result; S240, using the loss function to determine whether the training end condition is met, if not, adjusting the weight of the target detection model, and returning to S220 until the training end condition is met, to obtain a trained target detection model.

6. The method for detecting worker helmets in electric power construction scenarios according to claim 5 is characterized in that: S220 includes: S221, selecting an input image from the training set and inputting it into the backbone network, so as to extract features of the input image using the backbone network, and obtaining an input feature map X with a size of c×h×w; S222, the input feature map X of size c×h×w is processed by the first CBS module to obtain a feature map of size Feature map X1; S223, the size is The feature map X1 is processed by the second CBS module to produce a size of The feature map F; S224, the The feature map F is divided into four parts, which are convolved through four scale convolution channels to obtain a size of Four feature maps F0~F3; S225, the size is The four feature maps F0~F3 of are processed by the first splicing module to generate a multi-scale channel feature with a size of Feature map K1; S225, weighting the feature map K1 through the CA attention mechanism module to generate a refined feature map K2 with attention weights; S226, adding the refined feature map K2 to the feature map X1 to obtain the feature map X2, and outputting the feature map X2 and the feature map X1 to a feature map Y with a size of c×h×w through a fourth CBS module; S227, passing the feature map Y through the neck network and the output end to obtain a prediction result of the input image.

7. The method for detecting worker helmets in electric power construction scenes according to claim 6 is characterized in that: S225 includes: S2251, the CA attention mechanism module encodes the feature map K1 in the height and width directions respectively to generate a feature map in the height direction and a feature map in the width direction; S2252, concatenate the feature maps in the height direction and the width direction, perform 1×1 convolution, and activate them using the Sigmoid function to obtain an intermediate feature map f of size c×(h+w) containing width encoding and height encoding information; S2253, split the intermediate feature map f to generate two independent tensors f h ∈R C×H and f w ∈R C×W ; S2254, unify two independent tensors f using two 1×1 convolution operations respectively h ∈R C×H and f w ∈R C×w The number of channels is combined with the sigmoid activation function to obtain the attention vector; S2255, using the attention vector to perform attention weighting on the feature map K1 to obtain refined features on each channel; S2256, the refined features on each channel are combined into a refined feature map K2.

8. The method for detecting worker helmets in electric power construction scenes according to claim 7 is characterized in that: The feature map in the height direction and the feature map in the width direction in S2251 are expressed as: Where: z h Represents the feature map in the height direction; z w Represents the feature map in the width direction; Indicates the encoding value output by the cth channel in the height direction; Indicates the encoding value output by the cth channel in the width direction; w, h - the width and height of the image; x c (h, i) represents the value of feature map K1 at point (h, i) in the cth channel; The intermediate feature map f of S2252 is expressed as: f=δ(conv 1×1 ([z h ,z w ])) The specific formula of the Sigmoid activation function is as follows: The attention vector in S2254 is expressed as: The refined features on each channel in S2255 are expressed as: in, is the refined feature on the cth channel, is the feature map K1 on the cth channel, is the height-direction attention vector on the c-th channel, The width-wise attention vector on the c-th channel.

9. The method for detecting worker helmets in electric power construction scenes according to claim 6, characterized in that: The backbone network includes a CBS module, a CSP module and an SPPF module that introduce Ghost convolution, and the CBS module is a combination of a convolution module, a batch normalization module and a Sigmoid activation function; The CBS module introducing the Ghost convolution uses the Ghost convolution method to convolve the input image, and the process includes: Perform conventional convolution on the features of the previous layer to obtain the intrinsic feature map; Use a 3×3 convolution kernel to convolve the remaining features to obtain the Ghost feature map; The intrinsic feature map and the Ghost feature map are spliced ​​to obtain a CBS module spliced ​​feature map.

10. The method for detecting worker helmets in electric power construction scenarios according to claim 6, characterized in that: The loss function includes the confidence regression loss L obj , category loss L cls and the localization loss L loc ; The confidence regression loss is calculated using binary cross entropy and is expressed as: In the formula, λ obj , noobj Correspondingly, the coefficients of confidence regression error with and without targets; S- represents the number of grids used to divide the image in the backbone network; B represents the number of prediction boxes generated for each grid in the feature map; Indicates whether there is an object in the prediction box of the ith network. When the intersection and union ratio of the annotation box and the prediction box is greater than the set threshold, set is 1, otherwise it is 0; C i Indicates the confidence of the prediction box; Indicates the confidence of the annotation box; The positioning loss is calculated using CIOU, which is expressed as follows: Where b gt represents the coordinates of the center point of the annotation box; b represents the coordinates of the center point of the prediction box; c represents the diagonal length of the smallest circumscribed rectangle of the annotation box and the prediction box; ρ represents the Euclidean distance between the center points of the annotation box and the prediction box; α represents the weight parameter; v represents the aspect ratio of the prediction box, w gt Indicates the height of the annotation box, h gt represents the height of the annotation box, w represents the width of the prediction box, h represents the height of the prediction box; αv represents the penalty item set, the purpose of which is to make the width and height of the prediction box approach the annotation box faster.

Citation Information

Cited By

  • Vision-based chip placement system and operation method thereof

    CN120507640A