Object recognition method based on single-step multi-frame target detection (SSD) algorithm
By integrating a mixed attention mechanism and knowledge distillation into the SSD model, the system addresses resource allocation inefficiencies and training speed issues, improving object detection accuracy and speed in complex construction site environments.
Patent Information
- Application Number
- CN202510548332.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-15
AI Technical Summary
The existing SSD models cannot dynamically adjust resource allocation in the supervision of construction status of construction sites, and it is difficult to effectively identify the construction status in complex scenarios such as occlusion, uneven lighting, dense targets, and small targets, and the training speed is slow.
A single-step multi-frame object detection (SSD) algorithm is used to introduce a mixed attention mechanism and knowledge distillation technology, and VGG-16 is used as the basic network structure, combined with MobileNet v3 for model lightweighting and feature extraction, and the student model is trained through the teacher model to improve recognition accuracy and speed.
In complex construction scenarios, the recognition accuracy of small targets is improved, the allocation of resources to important areas is dynamically adjusted, the impact of background noise is reduced, and the recognition speed and accuracy of the model is improved.
Smart Images

Figure CN120318655A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of object recognition based on deep learning, and specifically to an image recognition method that introduces a hybrid attention mechanism into the teacher model SSD and uses MoblieNet v3 as the student model. Background Art
[0002] With the acceleration of the pace of urban construction, there are many new construction site projects, and the construction period is long, resulting in an increasing number of projects under construction at the same time. Therefore, the supervision of the construction status of construction sites has become particularly important. At present, the supervision of project illegal construction status mainly relies on two methods: manual online video viewing and offline patrol. These two supervision methods consume a lot of manpower and material resources and have low work efficiency.
[0003] With the continuous exploration and research of the concept of deep learning by researchers, and the continuous maturity of image recognition technology based on deep learning, researchers have applied recognition technology to the supervision of the construction status of construction sites. In the prior art, the VGG model of the convolutional neural network can be used. Through digital image processing technology, the images collected at the construction site are subjected to sample data enhancement and image smoothing preprocessing. Through the object detection algorithm and transfer learning method, the recognition speed of wearing safety belts and safety helmets is improved, the work efficiency of the supervisors of the construction status of construction sites is increased, and at the same time, the recognition accuracy is effectively improved. However, in the prior art, SSD pays the same attention to all regions and cannot dynamically adjust the resource allocation to more important target regions. Targets in complex construction scenarios are difficult to be effectively recognized and judged under the conditions of being blocked, uneven illumination, dense targets, small targets, etc. Moreover, the SSD dataset scale is complex and the training speed is slow.
[0004] To solve the above problems, the present application introduces an attention mechanism to enable SSD to dynamically adjust the resource allocation to more important target regions, effectively improving the recognition accuracy of small targets in complex construction scenarios of construction sites. At the same time, knowledge distillation is introduced to realize the lightweight of the model and improve the training speed of the SSD model. Summary of the Invention
[0005] The present invention uses the single-shot multibox detector (SSD) algorithm, takes VGG-16 as the basic network structure, introduces a hybrid attention mechanism, and adds knowledge distillation to the model. The student model is trained by the teacher model to realize the lightweight of the model, and at the same time improve the recognition accuracy mAP and recognition speed, effectively solving the disadvantages of low recognition accuracy and slow speed in traditional construction supervision scenarios.
[0006] The present invention provides an object recognition method based on the single-shot multibox detector (SSD) algorithm, including the following steps:
[0007] Step 1: Data Processing. Collect data from public datasets and actual scenarios, perform data cleaning and annotation. In terms of image enhancement, methods such as random cropping, scaling, flipping, modifying brightness, and adding noise are used to generate similar but different training samples, thereby expanding the scale of the training set. In addition, randomly changing the training samples reduces the model's dependence on certain attributes, thereby improving the model's generalization ability. Subsequently, the dataset is divided into a training set, a validation set, and a test set;
[0008] Step 2: Adjust the teacher model
[0009] 2.1 Input Preprocessing. Uniformly scale the input image to a fixed size of 300*300 and standardize the image quality;
[0010] 2.2 Feature Extraction
[0011] 2.2.1 Basic Network: Use the pre-trained convolutional network VGG-16 as the backbone network, adopt the first 13 convolutional layers of VGG16, and then replace the original FC layer with Conv 6 and Conv7 to extract basic features;
[0012] 2.2.2 Additional Convolutional Layers: SSD adds 4 deep convolutional network layers after the backbone network to extract higher-level semantic information. The number of channels of Conv8 is 512, and the number of channels of Conv9, Conv10, and Conv11 are all 256. From Conv7 to Conv11, the sizes of the output feature maps after these 5 convolutions are 19×19, 10×10, 5×5, 3×3, and 1×1 in sequence;
[0013] 2.2.3 Introduce the Convolutional Block Attention Module (CBAM): Insert the CBAM module into the backbone and additional convolutional layers of the SSD network. CBAM combines channel attention and spatial attention, and its working process is as follows: 1) Calculate channel attention to generate channel weights; 2) Calculate spatial attention to generate spatial weights; 3) Weight the feature map. The statements self.cbam1 = CBAM(512) and self.cbam2 = CBAM(1024) are used to insert CBAM at key layers and introduce it in the convolutional layer in the following way:
[0014]
[0015] 2.2.4 Bounding Box Feature Extraction Network (Multi-box Layers): SSD has a total of six multi-scale extraction networks (the feature maps obtained from the 4th, 7th, 8th, 9th, 10th, and 11th convolutional layers). Each layer performs convolution on loc and conf respectively to obtain the corresponding output;
[0016] Step 3: Adapt the student model structure
[0017] Use MobileNet v3 as the backbone network and make the following modifications:
[0018] 3.1 Add detection heads: Remove the original classification layer (Global Avg Pooling + FC layer), select the feature maps of the feature layers with stride = 16 and stride = 32 for output, add SSD-style detection heads, and adjust the output channels. Each detection head needs to output (num_anchors * (num_classes + 4)) channels to match the prediction format of SSD;
[0019] 3.2 Feature map channel alignment: Ensure the compatibility of the feature map channels between the student and teacher models: Adjust the feature map resolution through the stride to ensure the alignment of multi-scale features with SSD;
[0020] 3.3 Activation function and normalization adaptation: Since MobileNet-v3 uses h-swish while SSD commonly uses ReLU, use the following statements to unify the activation function:
[0021]
[0022] 3.4 Input allocation rate adaptation: Unify the input size. Since SSD uses 300 * 300 input, adjust the input layer of MobileNet-v3:
[0023] # Modify the first convolutional layer of MobileNet-v3 (originally adapted for 224x224)
[0024] self.features[0][0] = nn.Conv2d(3, 16, kernel_size = 3, stride = 2, padding = 1);
[0025] Step 4: Perform knowledge distillation
[0026] 4.1 Design the loss function
[0027] Total loss: Combine the supervision signal and the distillation loss: L student = γL detect + λ(L soft + L feat ), where L detect is the detection loss of the student model itself;
[0028] 4.2 Training
[0029] 4.2.1 Forward propagation
[0030] Teacher model: Feed the input image into the teacher model (SSD) to obtain the output of the teacher model, including the class probability distribution of the target and the position information of the bounding box;
[0031] Student model: Feed the same input image into the student model (MobileNet-v3) to obtain the output of the student model;
[0032] 4.2.2 Calculate the loss
[0033] According to the defined loss function, calculate the distillation loss, real loss, and total loss;
[0034] 4.2.3 Backpropagation
[0035] Through the backpropagation algorithm, calculate the gradient of the total loss with respect to the parameters of the student model. Use the Adam optimizer to update the parameters of the student model to continuously reduce the total loss;
[0036] 4.2.4 Iterative training
[0037] Repeat the above processes of forward propagation, calculating the loss, and backpropagation until the student model converges. During the training process, adjust the hyperparameters such as the learning rate and weight coefficient according to the performance of the validation set;
[0038] 4.3 Model evaluation
[0039] After the training is completed, use the test set to evaluate the student model. The evaluation metrics include the mean average precision (mAP), recall rate, accuracy, etc. Evaluate the effect of knowledge distillation by comparing with the performance of the teacher model and other baseline models;
[0040] 5. Image recognition
[0041] 5.1 Image input: Input the image to be recognized into the trained student model (MobileNet-v3);
[0042] 5.2 Feature extraction: The student model extracts features from the input image to obtain feature maps of different scales;
[0043] 5.3 Object detection: Perform object detection on the feature map to predict the class and position of the object;
[0044] 5.4 Post-processing: Perform post-processing on the prediction results, such as non-maximum suppression (NMS), to remove overlapping detection boxes and obtain the final detection results.
[0045] The object recognition method of the present invention has the following advantages:
[0046] 1. By introducing the attention mechanism, SSD dynamically adjusts resource allocation to more important target areas: 1) Dynamic weight allocation: After introducing the attention mechanism, SSD can adaptively learn which areas are more worthy of attention and assign higher weights, enabling the model to focus more on important objects in complex scenarios and reducing the influence of background noise; 2) Multi-level feature enhancement: Allows each layer to automatically adjust its contribution according to actual needs; 3) Without significantly increasing the computational cost, after introducing the attention mechanism, it has better mAP (mean Average Precision) performance indicators. Effective focus guidance makes the entire framework more sensitive and robust - even in the face of severely occluded or frequently pose-changing objects, it can maintain a high level of confidence scores.
[0047] 2. Introduce knowledge distillation to achieve model lightweight and improve the training speed of the SSD model: 1) Feature representation enhancement: The teacher model is usually trained on a large-scale dataset, and the learned feature representation is more abundant and robust. Through knowledge distillation, the student model can inherit these high-quality feature representations. This inheritance enables the student model to obtain better generalization ability even under limited data conditions. 2) Soft label guidance: Knowledge distillation not only transmits hard labels (i.e., the final classification results), but also uses soft labels in the form of probability distributions generated by the teacher model for more fine-grained learning. This approach helps reduce the risk of overfitting and improve the prediction accuracy in boundary cases. 3) Intermediate layer feature transfer: In some implementations, in addition to the output layer, feature maps are also extracted from different layers of the teacher network and mapped to the corresponding parts of the student network for matching. This method further promotes the understanding and reproduction ability of complex patterns. Brief Description of the Drawings
[0048] Figure 1 It is a flowchart of an object recognition method based on the single-shot multibox detector (SSD) algorithm of the present invention. Detailed Implementation Modes
[0049] The following further elaborates on the present invention in combination with embodiments and attached Figure 1 drawings.
[0050] Step 1: Data processing. Collect data from public datasets and actual scenarios, perform data cleaning and annotation. In terms of image enhancement, methods such as random cropping, scaling, flipping, modifying brightness, and adding noise are used to generate similar but different training samples, thereby expanding the scale of the training set. In addition, randomly changing the training samples reduces the model's dependence on certain attributes, thereby improving the model's generalization ability. Subsequently, divide the dataset into a training set, a validation set, and a test set;
[0051] Step 2: Adjust the teacher model
[0052] 2.1 Input preprocessing: uniformly scale the input image to a fixed size of 300*300 and standardize the image quality.
[0053] 2.2 Feature extraction
[0054] 2.2.1 Basic network: Use the pre-trained convolutional network VGG-16 as the backbone network, adopt the first 13 convolutional layers of VGG16, and then replace the original FC layer with Conv 6 and Conv7 to extract basic features.
[0055] 2.2.2 Additional convolutional layers: SSD adds 4 deep convolutional network layers after the backbone network to extract higher-level semantic information. The number of channels of Conv8 is 512, and the number of channels of Conv9, Conv10, and Conv11 are all 256. From Conv7 to Conv11, the sizes of the output feature maps after these 5 convolutions are 19×19, 10×10, 5×5, 3×3, and 1×1 in sequence.
[0056] 2.2.3 Introduce the hybrid attention mechanism: Insert the Convolutional Block Attention Module (CBAM) module into the backbone and additional convolutional layers of the SSD network. CBAM combines channel attention and spatial attention, and its working process is as follows: 1) Calculate channel attention to generate channel weights; 2) Calculate spatial attention to generate spatial weights; 3) Weight the feature map. The statements self.cbam1 = CBAM(512) and self.cbam2 = CBAM(1024) are used to insert CBAM at the key layers and introduce it in the convolutional layer in the following way:
[0057]
[0058] 2.2.4 Bounding box feature extraction network (Multi-box Layers): SSD has a total of six multi-scale extraction networks (feature maps obtained from the 4th, 7th, 8th, 9th, 10th, and 11th convolutional layers). Each layer performs convolution on loc and conf respectively to obtain the corresponding outputs.
[0059] Step 3: Adapt the student model structure
[0060] Use MobileNetv3 as the backbone network and make the following modifications:
[0061] 3.1 Adding a detection head: Remove the original classification layer (Global Avg Pooling + FC layer), select the feature maps output of the feature layers with stride = 16 and stride = 32, add an SSD-style detection head, and adjust the output channels. Each detection head needs to output (num_anchors * (num_classes + 4)) channels to match the prediction format of SSD;
[0062] 3.2 Feature map channel alignment: Ensure the compatibility of the feature map channels between the student and teacher models: Adjust the feature map resolution by the stride to ensure the alignment of multi-scale features with SSD;
[0063] 3.3 Activation function and normalization adaptation: Since MobileNet-v3 uses h-swish while SSD commonly uses ReLU, use the following statements to unify the activation function:
[0064] # Force the use of ReLU in the detection head
[0065] self.detection_head = nn.Sequential(
[0066] nn.Conv2d(...),
[0067] nn.ReLU(inplace = True),
[0068] nn.Conv2d(...) )
[0070] 3.4 Input allocation rate adaptation: Unify the input size. Since SSD uses 300 * 300 input, adjust the input layer of MobileNet-v3:
[0071] # Modify the first convolutional layer of MobileNet-v3 (originally adapted to 224x224)
[0072] self.features[0][0] = nn.Conv2d(3, 16, kernel_size = 3, stride = 2, padding = 1);
[0073] Step 4: Knowledge distillation
[0074] 4.1 Design the loss function
[0075] Total loss: Combine the supervision signal and the distillation loss: L student = γL detect + λ(L soft + L feat )), where L detect is the detection loss of the student model itself;
[0076] 4.2 Training
[0077] 4.2.1 Forward Propagation
[0078] Teacher model: Feed the input image into the teacher model (SSD) to obtain the output of the teacher model, including the class probability distribution of the target and the location information of the bounding box;
[0079] Student model: Feed the same input image into the student model (MobileNet-v3) to obtain the output of the student model;
[0080] 4.2.2 Calculate Loss
[0081] Calculate the distillation loss, real loss, and total loss according to the defined loss function;
[0082] 4.2.3 Backward Propagation
[0083] Through the backpropagation algorithm, calculate the gradient of the total loss with respect to the parameters of the student model. Use the optimizer Adam to update the parameters of the student model to continuously reduce the total loss;
[0084] 4.2.4 Iterative Training
[0085] Repeat the above processes of forward propagation, calculate loss, and backward propagation until the student model converges. During the training process, adjust the hyperparameters such as the learning rate and weight coefficient according to the performance of the validation set;
[0086] 4.2.5 Model Evaluation
[0087] After the training is completed, use the test set to evaluate the student model. The evaluation metrics include the mean average precision (mAP), recall rate, accuracy, etc. Evaluate the effect of knowledge distillation by comparing with the performance of the teacher model and other baseline models;
[0088] 5. Image Recognition
[0089] 1. Image Input: Input the image to be recognized into the trained student model (MobileNet-v3);
[0090] 2. Feature Extraction: The student model extracts features from the input image to obtain feature maps of different scales;
[0091] 3. Object Detection: Perform object detection on the feature map to predict the class and location of the object;
[0092] 4. Post-processing: Perform post-processing on the prediction results, such as non-maximum suppression (NMS), to remove overlapping detection boxes and obtain the final detection results.
[0093] The content not described in detail in this specification belongs to the prior art well-known to those of ordinary skill in the art.
Claims
1. An object recognition method based on the single-shot multibox detector (SSD) algorithm, characterized in that It includes the following steps: Step 1: Data processing. Collect data from public datasets and actual scenarios, perform data cleaning and annotation, and generate similar but different training samples by methods such as random cropping, scaling, flipping, modifying brightness, and adding noise to expand the scale of the training set. Divide the dataset into a training set, a validation set, and a test set; Step 2: Adjust the teacher model 2.1 Input preprocessing 2.2 Feature extraction 2.2.1 Construct the basic network: Use the pre-trained convolutional network VGG-16 as the backbone network, adopt the first 13 convolutional layers of VGG16, replace the original FC layer with Conv 6 and Conv7, and extract basic features; 2.2.2 Construct additional convolutional layers: SSD adds 4 deep convolutional network layers after the backbone network. The number of channels of Conv8 is 512, and the number of channels of Conv9, Conv10, and Conv11 are all 256. From Conv7 to Conv11, the sizes of the output feature maps after 5 convolutions are 19×19, 10×10, 5×5, 3×3, and 1×1 in sequence; 2.2.3 Introduce the hybrid attention mechanism: Insert the hybrid attention mechanism (CBAM) module into the backbone and additional convolutional layers of the SSD network; 2.2.4 Bounding box feature extraction network (Multi-box Layers): SSD has a total of six multi-scale extraction networks (feature maps obtained from the 4th, 7th, 8th, 9th, 10th, and 11th convolutional layers). Each layer performs convolution on loc and conf respectively to obtain the corresponding output; Step 3: Adapt the student model structure Use MobileNet v3 as the backbone network and make the following modifications: 3.1 Add detection heads: Remove the original classification layer (GlobalAvg Pooling+FC layer), select the feature maps of the feature layers with stride = 16 and stride = 32 for output, add SSD-style detection heads, and adjust the output channels. Each detection head outputs (num_anchors*(num_classes+4)) channels to match the prediction format of SSD; 3.2 Feature map channel alignment: Ensure the compatibility of the feature map channels between the student and teacher models: Adjust the feature map resolution through the stride to ensure the alignment of multi-scale features with SSD; 3.3 Activation function and normalization adaptation 3.4 Input allocation rate adaptation Step 4: Perform knowledge distillation 4.1 Design the loss function Total loss: combined supervision signal and distillation loss: L student = γL detect + λ(L soft + L feat ), where L detect is the detection loss of the student model itself; 4.2 Training 4.2.1 Forward propagation Teacher model: Send the input image into the teacher model (SSD) to obtain the output of the teacher model, including the class probability distribution of the target and the position information of the bounding box; Student model: Send the same input image into the student model (MobileNet-v3) to obtain the output of the student model; 4.2.2 Calculate the loss According to the defined loss function, calculate the distillation loss, the real loss, and the total loss; 4.2.3 Backward propagation Through the backpropagation algorithm, calculate the gradient of the total loss with respect to the parameters of the student model, and use the optimizer Adam to update the parameters of the student model to continuously reduce the total loss; 4.2.4 Iterative Training Repeat the above processes of forward propagation, loss calculation, and backpropagation until the student model converges. During the training process, adjust hyperparameters such as the learning rate and weight coefficients according to the performance on the validation set; 4.3 Model Evaluation After training is completed, use the test set to evaluate the student model. The evaluation metrics include the mean average precision (mAP), recall rate, accuracy, etc. Evaluate the effect of knowledge distillation by comparing with the performance of the teacher model and other baseline models; 5. Image Recognition 5.1 Image Input: Input the image to be recognized into the trained student model (MobileNet-v3); 5.2 Feature Extraction: The student model extracts features from the input image to obtain feature maps of different scales; 5.3 Object Detection: Perform object detection on the feature map to predict the category and location of the object; 5.4 Post-Processing: Perform post-processing on the prediction results, such as non-maximum suppression (NMS), to remove overlapping detection boxes and obtain the final detection results.
2. The recognition method according to claim 1, wherein: In step 2.1, the input preprocessing is to uniformly scale the input image to a fixed size of 300*300 and standardize the image quality.
3. The recognition method according to claim 1, characterized in that: In step 2.2.3, the Convolutional Block Attention Module (CBAM) consists of two modules, the Channel Attention Module (CAM) and the Spatial Attention Module (SAM), combined in a serial manner. The specific working process is as follows: 1) Calculate the channel attention to generate channel weights; 2) Calculate the spatial attention to generate spatial weights; 3) Weight the feature map; Use the statements self.cbam1 = CBAM(512) and self.cbam2 = CBAM(1024) to insert CBAM at the key layers and introduce it in the convolutional layer in the following way: fork in range(25): x = self.model[k](x) x = self.cbam1(x) # Add CBAM sources.append(self.L2Norm(x)).
4. The recognition method according to claim 1, characterized in that: In step 3.3, use the following statements to unify the activation function: # Force the use of ReLU in the detection head self.detection_head = nn.Sequential( nn.Conv2d(...), nn.ReLU(inplace=True), nn.Conv2d(...) )。 5. The recognition method according to claim 2, characterized in that: In step 3.4, unify the input size and adjust the input layer of MobileNet-v3: # Modify the first convolutional layer of MobileNet-v3 (originally adapted for 224x224) self.features[0][0] = nn.Conv2d(3, 16, kernel_size = 3, stride = 2, padding = 1).