Target detection model system and method based on lightweight network

Through the combination of lightweight network and feature pyramid module, the detection accuracy and multi-scale adaptability of the object detection model on resource-constrained devices are solved, and efficient and stable object detection is achieved, suitable for edge computing and the Internet of Things.

CN120355936AInactive Publication Date: 2025-07-22GUANGXI POLICE ACAD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510264008.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing object detection model is difficult to ensure detection accuracy while reducing the computational complexity, and there is insufficient optimization in multi-scale feature fusion, making it difficult to realize real-time application on resource-constrained devices.

Method used

The lightweight network structure is adopted, combined with the depth-separable convolution and feature pyramid module, image preprocessing, feature extraction, multi-scale feature fusion, target classification and positioning, and the detection results are optimized through non-maximum suppression method.

Benefits of technology

While maintaining detection accuracy, the model's operating efficiency on low-power devices is improved. It is suitable for scenarios with high real-time requirements. It can stably detect targets of different sizes, improve detection accuracy and accuracy, and is suitable for edge computing and IoT environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355936A_ABST
    Figure CN120355936A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection model system and method based on a lightweight network, and belongs to the technical field of target detection, and the system comprises an input module which is used for the reading and preprocessing of an input image, so as to optimize the format and quantity of input data of a model; the feature extraction module is used for extracting key features of the image by using a lightweight neural network; the feature pyramid module is used for fusing features on different scales to adapt to multi-scale target detection; according to the invention, the multi-scale feature fusion design enables the system to be more stable when processing a small target and a large target, and the detection precision is higher; flexible target classification and positioning: a detection head module in the system is designed with classification and regression branches, efficient target classification and bounding box prediction can be carried out according to features extracted by a feature pyramid, and in combination with a loss function based on cross entropy and smooth L1 loss, the system can effectively reduce classification errors and bounding box regression errors, so that the accuracy of target classification and positioning is improved. The detection precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of object detection, and particularly relates to an object detection model system and method based on a lightweight network. Background Art

[0002] Object detection is a core task in computer vision and is widely applied in fields such as autonomous driving, security monitoring, and face recognition. Traditional object detection algorithms are mainly based on convolutional neural networks, such as Faster R-CNN, SSD, YOLO, etc. These models perform excellently in feature extraction, classification, and regression tasks. However, traditional detection models usually have complex calculations and a large number of parameters, making it difficult to achieve real-time applications on resource-constrained devices (such as embedded devices and mobile terminals). This limitation significantly affects the application of object detection technology in scenarios such as edge computing and the Internet of Things.

[0003] To address this challenge, the design of lightweight networks has become a research hotspot in recent years. Lightweight networks improve efficiency by reducing the amount of computation and the number of parameters and are suitable for deployment on low-power devices. These networks usually adopt techniques such as depthwise separable convolutions and group convolutions to reduce the model complexity, significantly reducing the computational cost while ensuring accuracy. Object detection models based on lightweight networks combine the advantages of efficient feature extraction and fast inference and are gradually applied in scenarios such as real-time detection on mobile terminals.

[0004] However, current lightweight object detection models still face some technical problems: how to ensure detection accuracy while reducing model complexity, and how to optimize multi-scale feature fusion to meet the detection requirements of targets of different sizes.

[0005] Based on this, the present invention designs an object detection model system and method based on a lightweight network to solve the above problems. Summary of the Invention

[0006] The purpose of the present invention is to propose an object detection model system and method based on a lightweight network to solve the problems of how to ensure detection accuracy while reducing model complexity and how to optimize multi-scale feature fusion to meet the detection requirements of targets of different sizes.

[0007] To achieve the above purpose, the present invention adopts the following technical solutions:

[0008] An object detection model system based on a lightweight network, comprising:

[0009] An input module: responsible for reading and preprocessing the input image to optimize the input data format and quantity of the model;

[0010] Feature extraction module: used to extract key features of an image using a lightweight neural network;

[0011] Feature pyramid module: used to fuse features at different scales to adapt to multi-scale object detection;

[0012] Detection head module: used to classify and locate objects for the features extracted by the feature pyramid;

[0013] Loss calculation module: used to calculate the losses of classification and regression to guide network training;

[0014] Post-processing module: used to process the output results of the model by the non-maximum suppression (NMS) method and output the final detection boxes.

[0015] As a further description of the above technical solution: the preprocessing of the image includes image scaling, normalization, and data augmentation;

[0016] Image scaling: scaling the input image to the input size required by the model, usually 224×224, 320×320 or other sizes, to meet the requirements of the network and reduce the computational burden;

[0017] Normalization processing: mapping the pixel values from the original range of 0 - 255 to between [0,1] or [-1,1] to eliminate the differences in different image distributions and improve the generalization ability of the model;

[0018] The formula is:

[0019] where I is the pixel value of the original image, and μ and σ are the mean and standard deviation of the image respectively;

[0020] Data augmentation: operations such as random cropping, horizontal flipping, and brightness adjustment, enabling the model to have stronger generalization ability, enabling the model to handle different shooting angles and lighting changes, and improving the detection accuracy of objects.

[0021] As a further description of the above technical solution: the specific steps for extracting the key features of the image are as follows:

[0022] Convolutional layer extracts low-level features: used to initially extract low-level features of the image;

[0023] The formula used is:

[0024] where X represents the input feature map, with dimensions H×W×M (height, width, and number of channels);

[0025] K represents the convolutional kernel, with dimensions h×w×M×N (height, width, number of input channels, and number of output channels) of the convolutional kernel;

[0026] Y is the output feature map with dimensions H′×W′×N;

[0027] i and j are the position indices of the output feature map, and k is the output channel index;

[0028] This formula indicates that the standard convolution performs a pixel-by-pixel dot product operation on the input feature map X using the convolution kernel K to generate the output feature map Y;

[0029] Depthwise separable convolution: To reduce the computational complexity and the number of parameters, depthwise separable convolution is used to replace the standard convolution. This convolution is divided into two steps:

[0030] Depthwise convolution: Perform convolution operations separately on each channel to extract the features of each channel individually;

[0031]

[0032] Among them, X represents the input feature map with dimensions H×W×M;

[0033] K represents the convolution kernel with dimensions h×w×M (i.e., there is an independent convolution kernel for each input channel);

[0034] Y is the output feature map with dimensions H′×W′×M;

[0035] i and j are the position indices of the output feature map, and m represents the channel index;

[0036] Pointwise convolution: Use 1×1 convolution to integrate the features of each channel to obtain the fused features; Depthwise separable convolution has a relatively small computational amount and is suitable for lightweight network applications;

[0037]

[0038] Among them, X represents the input feature map with dimensions H×W×M;

[0039] W is the weight of the pointwise convolution with dimensions M×N;

[0040] Y is the output feature map with dimensions H×W×N;

[0041] i and j are the spatial position indices of the output feature map, and m and n are the input channel index and the output channel index respectively;

[0042] Activation and pooling: After each convolutional layer, use activation functions and pooling operations (such as max pooling) to further extract and reduce the size of the feature map, thereby improving the computational efficiency and increasing the non-linear expression ability of the network;

[0043] The activation function is:

[0044] To further reduce the size of the feature map, a pooling operation is used. The specific formula is as follows:

[0045]

[0046] Where X is the input feature map, Y is the output feature map after pooling. The maximum pooling takes the maximum value within a specific window, reducing the spatial dimension and retaining the most significant feature information.

[0047] As a further description of the above technical solution: The multi-scale fused features mainly include the following steps:

[0048] Construct multi-scale feature layers: Select feature maps at different levels from the feature extraction module to generate multiple feature layers with different scales to contain rich hierarchical information;

[0049] Top-down feature fusion: Adopt a top-down approach to gradually downsample and fuse high-level semantic features into low-level feature layers to ensure that the feature map contains both high-level semantic information and fine-grained spatial information;

[0050] Upsampling and merging: For the high-level features of each layer, perform upsampling operations and add them to the corresponding low-level features. After merging, generate a feature map with a pyramid structure;

[0051] The formula is, F out = F top + Upsample(F bottom );

[0052] Where F out represents the feature of the previous level, F bottom represents the feature of the current level. After the merging operation, more rich multi-scale features are generated.

[0053] As a further description of the above technical solution: The main steps for target classification and localization include:

[0054] Anchor box generation: Generate multiple anchor boxes with different sizes and aspect ratios at the grid points of each feature map;

[0055] Classification branch: For each anchor box, output a class probability through a convolutional layer. Usually, the softmax or sigmoid function is used to calculate the probability of each class;

[0056] Formula:

[0057] Where, y i is the true label, p i is the predicted probability;

[0058] Regression branch: Output the position offset of each anchor box relative to the ground truth box. The regression branch outputs the coordinate offsets (i.e., dx, dy, dw, dh) through convolutional layers, calculates the regression loss for each bounding box, and generally uses smooth L1 loss for the regression loss:

[0059]

[0060] Among them, t i is the predicted bounding box coordinate, and

[0061] As a further description of the above technical solution: The specific steps for calculating the classification loss and regression loss are as follows:

[0062] Calculate the classification loss: Use cross-entropy loss to measure the classification ability of the model;

[0063] Calculate the localization loss: Use smooth L1 loss to evaluate the matching degree between the predicted box and the ground truth box;

[0064] Combine the losses: Combine the classification and localization losses to generate the total loss function;

[0065] The formula is, L = λ cls ·L cls + λ loc ·L loc ;

[0066] Among them, λ cls and λ loc are the weights of the classification loss and the localization loss, used to balance the influence of the two.

[0067] As a further description of the above technical solution: The specific steps for processing the output result of the detection head to generate the final detection box are as follows:

[0068] Confidence screening: Screen out the candidate boxes with high confidence according to the classification probability of the detection head, and remove the detection boxes with low confidence;

[0069] Non-maximum suppression (NMS): Remove the overlapping candidate boxes by calculating the IoU (Intersection over Union), and keep the box with the highest score as the final detection result;

[0070] The formula is,

[0071] Among them, A and B are the candidate box regions;

[0072] Output the detection result: Take the detection box after NMS processing as the final output, including the category and location of the target.

[0073] As a further description of the above technical solution: The underlying features include the basic edges and texture information of the image.

[0074] As a further description of the above technical solution: The sizes of the anchor boxes cover the possible target scales to improve the capture ability for targets.

[0075] A target detection model method based on a lightweight network, the method comprising the following steps:

[0076] Reading and preprocessing the input image to optimize the input data format and quantity of the model;

[0077] Using a lightweight neural network to extract the key features of the image;

[0078] Fusing features at different scales to adapt to multi-scale target detection;

[0079] Performing target classification and localization on the features extracted by the feature pyramid;

[0080] Calculating the losses of classification and regression to guide network training;

[0081] Processing the output result of the model by the non-maximum suppression (NMS) method to output the final detection box.

[0082] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are as follows:

[0083] 1. In the present invention, the system adopts a lightweight network structure and uses depthwise separable convolution to reduce the number of parameters and the amount of calculation, greatly improving the operation efficiency of the model on low-power devices. This system can achieve efficient inference while maintaining detection accuracy and is suitable for scenarios with high real-time requirements; Multi-scale feature fusion: By introducing a Feature Pyramid Network (FPN), this system can fuse image features at different scales to enhance the detection effect for targets of different sizes. The top-down feature fusion strategy of the FPN ensures that the semantic information of high-level features can be transmitted to low-level features, and at the same time, the detailed information of low-level features can also be used for large target detection. This multi-scale feature fusion design makes the system more stable and has higher detection accuracy when dealing with small and large targets; Flexible target classification and localization: The detection head module in the system is designed with classification and regression branches, which can perform efficient target classification and bounding box prediction based on the features extracted by the feature pyramid. Combining with a loss function based on cross-entropy and smooth L1 loss, this system can effectively reduce classification errors and bounding box regression errors and improve detection accuracy.

[0084] 2. In the present invention, redundant candidate bounding boxes are removed through non-maximum suppression (NMS). The system can effectively improve the accuracy of the detection results. NMS combines the intersection over union (IoU) to screen candidate bounding boxes with high overlap degrees, ensuring that the finally output detection bounding boxes have high confidence and position accuracy.

[0085] 3. In the present invention, the lightweight design of the system makes it easy to be deployed in resource-constrained environments such as mobile devices and embedded systems, suitable for edge computing and Internet of Things scenarios. The modular design of the model structure facilitates flexible adjustment and expansion. Developers can optimize the parameter settings of each module according to the specific requirements of the application scenario, so as to achieve the best detection performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 It is a schematic diagram of the module composition of a target detection model system and method based on a lightweight network proposed by the present invention;

[0087] Figure 2 It is a schematic diagram of the steps of a target detection model system and method based on a lightweight network proposed by the present invention;

[0088] Figure 3 It is a schematic diagram of the feature pyramid process of a target detection model system and method based on a lightweight network proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0089] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0090] Please refer to the attached Figure 1 - attached Figure 3 , the present invention provides a technical solution: a target detection model system based on a lightweight network, including:

[0091] Input module: responsible for reading and preprocessing the input image to optimize the input data format and quantity of the model;

[0092] Feature extraction module: used to extract key features of the image by using a lightweight neural network;

[0093] Feature pyramid module: used to fuse features at different scales to adapt to multi-scale target detection;

[0094] Detection head module: used to perform target classification and localization on the features extracted by the feature pyramid;

[0095] Loss calculation module: used to calculate the losses of classification and regression to guide network training;

[0096] Post-processing module: used to process the output results of the model by the non-maximum suppression (NMS) method and output the final detection boxes.

[0097] The preprocessing of the image includes image scaling, normalization, and data augmentation;

[0098] Image scaling: Scale the input image to the input size required by the model, usually 224×224, 320×320 or other sizes, to meet the requirements of the network and reduce the computational burden;

[0099] Normalization processing: Map the pixel values from the original range of 0 - 255 to between [0,1] or [-1,1] to eliminate the differences in different image distributions and improve the generalization ability of the model;

[0100] The formula is:

[0101] where I is the pixel value of the original image, and μ and σ are the mean and standard deviation of the image respectively;

[0102] Data augmentation: Operations such as random cropping, horizontal flipping, and brightness adjustment are performed to make the model have stronger generalization ability. The model can handle different shooting angles and lighting changes, and improve the detection accuracy of the target.

[0103] Specifically, extracting the key features of the image includes the following steps:

[0104] The convolutional layer extracts low-level features: used to initially extract the low-level features of the image;

[0105] The formula adopted is:

[0106] where X represents the input feature map, with dimensions H×W×M (height, width, and number of channels);

[0107] K represents the convolutional kernel, with dimensions h×w×M×N (height, width, number of input channels, and number of output channels) of the convolutional kernel;

[0108] Y is the output feature map, with dimensions H′×W′×N;

[0109] i, j are the position indices of the output feature map, and k is the output channel index;

[0110] This formula means that the standard convolution performs a pixel-by-pixel dot product operation on the input feature map X with the convolutional kernel K to generate the output feature map Y;

[0111] Depthwise Separable Convolution: To reduce the computational complexity and the number of parameters, depthwise separable convolution is used to replace the standard convolution. This convolution is divided into two steps:

[0112] Depthwise Convolution: The convolution operation is performed separately on each channel to extract the features of each channel individually;

[0113]

[0114] Among them, X represents the input feature map, with dimensions of H×W×M;

[0115] K represents the convolution kernel, with dimensions of h×w×M (that is, each input channel has an independent convolution kernel);

[0116] Y is the output feature map, with dimensions of H′×W′×M;

[0117] i, j are the position indices of the output feature map, and m represents the channel index;

[0118] Pointwise Convolution: 1×1 convolution is used to integrate the features of each channel to obtain the fused features; Depthwise separable convolution has a relatively small computational amount and is suitable for lightweight network applications;

[0119]

[0120] Among them, X represents the input feature map, with dimensions of H×W×M;

[0121] W is the weight of the pointwise convolution, with dimensions of M×N;

[0122] Y is the output feature map, with dimensions of H×W×N;

[0123] i, j are the spatial position indices of the output feature map, and m and n are the input channel index and the output channel index respectively;

[0124] Activation and Pooling: After each convolutional layer, an activation function and a pooling operation (such as max pooling) are used to further extract and reduce the size of the feature map, thereby improving the computational efficiency and increasing the non-linear expression ability of the network;

[0125] The activation function is:

[0126] To further reduce the size of the feature map, a pooling operation is used. The specific formula is:

[0127]

[0128] Among them, X is the input feature map, Y is the output feature map after pooling, and max pooling takes the maximum value within a specific window, reducing the spatial dimension and retaining the most significant feature information.

[0129] The main steps for fusing features at multiple scales include the following:

[0130] Constructing multi-scale feature layers: Select feature maps at different levels from the feature extraction module to generate multiple feature layers with different scales to contain rich hierarchical information;

[0131] Top-down feature fusion: Adopt a top-down approach to gradually downsample and fuse high-level semantic features into low-level feature layers to ensure that the feature maps contain both high-level semantic information and fine-grained spatial information;

[0132] Upsampling and merging: For the high-level features of each layer, perform upsampling operations and add them to the corresponding low-level features. After merging, generate feature maps with a pyramid structure;

[0133] The formula is, F out = F top + Upsample(F bottom );

[0134] where F out represents the features of the previous level, and F bottom represents the features of the current level. After the merging operation, more rich multi-scale features are generated.

[0135] The main steps for object classification and localization include:

[0136] Anchor box generation: Generate multiple anchor boxes with different sizes and aspect ratios at the grid points of each feature map;

[0137] Classification branch: For each anchor box, output a class probability through a convolutional layer. Usually, the softmax or sigmoid function is used to calculate the probability of each class;

[0138] Formula:

[0139] where, y i is the true label, and p i is the predicted probability;

[0140] Regression branch: Output the position offset of each anchor box relative to the true box. The regression branch outputs the coordinate offsets (i.e., dx, dy, dw, dh) through a convolutional layer, calculates the regression loss of each bounding box, and the regression loss generally uses smooth L1 loss:

[0141]

[0142] where, t i is the predicted bounding box coordinate, is the true bounding box.

[0143] The specific steps for calculating the classification loss and the regression loss are as follows:

[0144] Calculate the classification loss: Use the cross-entropy loss to measure the classification ability of the model;

[0145] Calculate the localization loss: Use the smooth L1 loss to evaluate the matching degree between the predicted bounding box and the ground truth bounding box;

[0146] Combine the losses: Integrate the classification and localization losses to generate the total loss function;

[0147] The formula is, L = λ cls ·L cls +λ loc ·L loc ;

[0148] Where, λ cls and λ loc are the weights of the classification loss and the localization loss, used to balance the influence of the two.

[0149] Process the output results of the detection head to generate the final detection bounding box, which specifically includes the following steps:

[0150] Confidence screening: Screen out the candidate bounding boxes with high confidence according to the classification probability of the detection head, and remove the detection bounding boxes with low confidence;

[0151] Non-maximum suppression (NMS): Remove the overlapping candidate bounding boxes by calculating the IoU (Intersection over Union), and keep the bounding box with the highest score as the final detection result;

[0152] The formula is,

[0153] Where, A and B are the candidate bounding box regions;

[0154] Output the detection result: Take the detection bounding box after NMS processing as the final output, including the category and location of the target.

[0155] The underlying features include the basic edge and texture information of the image.

[0156] The sizes of the anchor boxes cover the possible target scales to improve the ability to capture targets.

[0157] A method for a target detection model based on a lightweight network, the method includes the following steps:

[0158] Read and preprocess the input image to optimize the input data format and quantity of the model;

[0159] Use the lightweight neural network to extract the key features of the image;

[0160] Fuse the features at different scales to adapt to multi-scale target detection;

[0161] Perform object classification and localization on the features extracted by the feature pyramid;

[0162] Calculate the losses for classification and regression to guide network training;

[0163] Process the output results of the model by the non-maximum suppression (NMS) method to output the final detection boxes.

[0164] As described above, it is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.

Claims

1. A target detection model system based on a lightweight network, characterized in that, It includes: Input module: Responsible for reading and preprocessing the input image to optimize the input data format and quantity of the model; Feature extraction module: Used to extract the key features of the image by using a lightweight neural network; Feature pyramid module: Used to fuse features at different scales to adapt to multi-scale object detection; Detection head module: Used to classify and locate objects for the features extracted by the feature pyramid; Loss calculation module: Used to calculate the losses of classification and regression to guide network training; Post-processing module: Used to process the output results of the model by the non-maximum suppression (NMS) method and output the final detection boxes.

2. The object detection model system based on a lightweight network according to claim 1, characterized in that, The preprocessing of the image includes image scaling, normalization, and data augmentation; Image scaling: Scale the input image to the input size required by the model, usually 224×224, 320×320 or other sizes, to meet the requirements of the network and reduce the computational burden; Normalization processing: Map the pixel values from the original range of 0 - 255 to between [0,1] or [-1,1] to eliminate the differences in different image distributions and improve the generalization ability of the model; The formula is as follows: where I is the pixel value of the original image, and μ and σ are the mean and standard deviation of the image respectively; Data augmentation: Operations such as random cropping, horizontal flipping, and brightness adjustment are performed to make the model have stronger generalization ability. The model can handle different shooting angles and lighting changes, and improve the detection accuracy of objects.

3. The object detection model system based on a lightweight network according to claim 1, wherein The specific steps for extracting the key features of the image are as follows: Convolutional layer extracts low-level features: Used to initially extract the low-level features of the image; The formula used is: where X represents the input feature map, with dimensions of H×W×M (height, width, and number of channels); K represents the convolutional kernel, with dimensions of h×w×M×N (height, width, number of input channels, and number of output channels of the convolutional kernel); Y is the output feature map, with dimensions of H′×W′×N; i, j are the position indices of the output feature map, and k is the output channel index; This formula indicates that the standard convolution performs a pixel-by-pixel dot product operation on the input feature map X with the convolutional kernel K to generate the output feature map Y; Depthwise separable convolution: To reduce the computational complexity and the number of parameters, depthwise separable convolution is used to replace the standard convolution. This convolution is divided into two steps: Depthwise convolution: Perform convolution operations separately on each channel to extract the features of each channel independently; where X represents the input feature map, with dimensions of H×W×M; K represents the convolutional kernel, with dimensions of h×w×M (i.e., each input channel has an independent convolutional kernel); Y is the output feature map, with dimensions of H′×W′×M; i, j are the position indices of the output feature map, and m represents the channel index; Pointwise convolution: Use 1×1 convolution to integrate the features of each channel to obtain the fused features; Depthwise separable convolution has a relatively small computational amount and is suitable for lightweight network applications; where X represents the input feature map, with dimensions of H×W×M; W is the weight of the pointwise convolution, with dimensions of M×N; Y is the output feature map, with dimensions of H×W×N; i, j are the spatial position indices of the output feature map, and m and n are the input channel index and output channel index respectively; Activation and Pooling: After each convolutional layer, activation functions and pooling operations (such as max pooling) are used to further extract and reduce the size of the feature maps, thereby improving computational efficiency and increasing the network's non-linear representation ability; The activation function is: To further reduce the size of the feature maps, pooling operations are used. The specific formula is: Where X is the input feature map, Y is the output feature map after pooling. Max pooling takes the maximum value within a specific window, reducing the spatial dimension and retaining the most significant feature information.

4. A target detection model system based on a lightweight network according to claim 1, characterized in that, The multi-scale feature fusion mainly includes the following steps: Constructing multi-scale feature layers: Select feature maps of different levels from the feature extraction module to generate multiple feature layers of different scales to contain rich hierarchical information; Top-down feature fusion: Adopting a top-down approach, the high-level semantic features are gradually downsampled and fused into the low-level feature layers to ensure that the feature maps contain both high-level semantic information and fine-grained spatial information; Upsampling and merging: For the high-level features of each layer, perform upsampling operations and add them to the corresponding low-level features. After merging, a feature map with a pyramid structure is generated; The formula is, F out = F top + Upsample(F bottom ); Among which F out represents the features of the upper level, and F bottom represents the features of the current layer, and richer multi-scale features are generated after the merging operation.

5. The object detection model system based on a lightweight network according to claim 1, wherein, The main steps for target classification and localization include: Anchor box generation: Generate multiple anchor boxes of different sizes and aspect ratios at the grid points of each feature map; Classification branch: For each anchor box, output a class probability through a convolutional layer. Usually, softmax or sigmoid functions are used to calculate the probability of each class; Formula: where y i is the true label, and p i is the predicted probability; Regression branch: Output the position offset of each anchor box relative to the ground truth box. The regression branch outputs the coordinate offsets (i.e., dx, dy, dw, dh) through a convolutional layer, calculates the regression loss of each bounding box, and the regression loss generally uses smooth L1 loss: where t i are the predicted bounding box coordinates, is the ground truth bounding box.

6. The object detection model system based on a lightweight network according to claim 1, wherein, The specific steps for calculating the classification loss and regression loss are as follows: Calculating the classification loss: Use cross-entropy loss to measure the classification ability of the model; Calculating the localization loss: Use smooth L1 loss to evaluate the matching degree between the predicted box and the ground truth box; Combined loss: Combine the classification and localization losses to generate the total loss function; The formula is L = λ cls ·L cls +λ loc ·L loc ; Among them, λ cls and λ loc are the weights of the classification loss and the localization loss, which are used to balance the influence of the two.

7. The object detection model system based on a lightweight network according to claim 1, characterized in that, Processing the output results of the detection head to generate the final detection box, which specifically includes the following steps: Confidence screening: Screen out the candidate boxes with high confidence according to the classification probability of the detection head, and remove the detection boxes with low confidence; Non-maximum suppression (NMS): Remove the overlapping candidate boxes by calculating the IoU (Intersection over Union), and retain the box with the highest score as the final detection result; The formula is Where A and B are the candidate box regions; Outputting the detection results: Use the detection box after NMS processing as the final output, including the class and location of the target.

8. The object detection model system based on a lightweight network according to claim 3, characterized in that, The low-level features include the basic edges and texture information of the image.

9. The object detection model system based on a lightweight network according to claim 5, wherein The sizes of the anchor boxes cover the possible target scales to improve the capture ability of the targets.

10. A method for an object detection model based on a lightweight network. According to the object detection model system based on a lightweight network described in any one of claims 1-9, it is characterized in that The method includes the following steps: Reading and preprocessing the input image to optimize the input data format and quantity of the model; Using a lightweight neural network to extract the key features of the image; Fusing features at different scales to adapt to multi-scale object detection; Performing target classification and localization on the features extracted from the feature pyramid; Calculating the losses of classification and regression to guide network training; Process the output results of the model by the non-maximum suppression (NMS) method to output the final detection boxes.

Citation Information

Cited By

  • Interactive refined detection method and device for unmanned aerial vehicle inspection image target detection

    CN121121572A

  • Stress injury automatic detection method based on YOLO neural network

    CN121190414A