Elevator Electric Bicycle Recognition Method Based on Improved YOLOv8
By building an improved electric vehicle recognition model based on YOLOv8, using local aggregation module and multi-dimensional collaborative attention module to enhance feature extraction and interaction, the problem of low recognition efficiency when electric vehicles enter the elevator is solved, and higher recognition accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202411687121.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2044-11-25
AI Technical Summary
The prior art has low recognition efficiency when electric vehicles enter the elevator, especially when small object detection, different angles, partial occlusion or extreme lighting conditions, and the model may have overfitting problems.
Based on YOLOv8, an improved electric vehicle recognition model is constructed, which consists of a feature extraction backbone network module, an aggregation neck network module, a multi-dimensional collaborative attention module and a candidate box prediction module. Feature extraction and feature interaction capabilities are enhanced through the local aggregation module and a multi-dimensional collaborative attention module.
It significantly improves the recognition accuracy and robustness of electric vehicles in complex scenarios, especially in the recognition of small electric vehicles and edge details, and improves the overall performance of the identification model.
Smart Images

Figure CN119181012B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and specifically to an elevator electric vehicle recognition method based on improved YOLOv8. Background Art
[0002] With the continuous advancement of urbanization construction, more and more residential buildings in communities have installed elevators. At the same time, the popularization of electric vehicles has brought convenience, but safety accidents caused by electric vehicles entering elevators occur frequently. After an electric vehicle is ridden, the battery and motor are at a relatively high temperature, while the space in the elevator is narrow and the air circulation is poor. When an electric vehicle enters the elevator, the vehicle is likely to collide with the elevator door or side wall, and such an impact is very likely to cause the electric vehicle to catch fire or even explode. Therefore, in order to prevent safety accidents caused by electric vehicles entering elevators, it is necessary to identify electric vehicles in the elevator and take preventive measures in advance to ensure the travel safety of residents.
[0003] Regarding the problem of electric vehicles entering elevators, current traditional identification means include physical vehicle blocking, weight sensor identification, and visual identification. Physical vehicle blocking is to manually set up a railing to block electric vehicles at the elevator entrance. However, this method not only restricts the entry of electric vehicles but also blocks other travel tools such as bicycles and children's bicycles, bringing inconvenience and being too costly; the weight sensor identifies electric vehicles by measuring the weight in the elevator, but in the case of a large number of passengers, this method may not be able to accurately determine whether an electric vehicle has entered. In contrast, the visual identification method uses computer vision technology and deep learning methods to capture images of electric vehicles through a camera and identify the characteristics of electric vehicles through image processing algorithms. This method has high accuracy and fast response speed and has great development potential, and is becoming the main means to solve the problem of electric vehicle identification in the future.
[0004] In current deep learning electric vehicle recognition methods, there are mainly detection algorithms based on CNN and models based on Transformer. However, the Transformer-based model cannot achieve real-time detection and requires excessive computing resources. In the CNN-based detection algorithms, algorithms like Faster R-CNN and SSD have low recognition efficiency in complex occlusion environments. Therefore, the YOLO framework based on CNN is the best choice for real-time electric vehicle detection. YOLO is a real-time detection framework widely used in computer vision object detection tasks and can complete object recognition in a single inference process. Compared with other detection frameworks, it has a faster recognition speed and is applicable to various scenarios including images, videos, and real-time streaming media. As the most practical version of this series, YOLOv8 further improves the detection speed and accuracy and provides a unified framework for model training. However, even so, the YOLOv8 framework also has its own limitations or drawbacks when recognizing electric vehicles. First, in terms of small object detection, YOLOv8 has insufficient detection accuracy for electric vehicles with small sizes. Second, the appearance of electric vehicles changes greatly at different angles, with partial occlusion, or under extreme lighting conditions, which causes the traditional YOLOv8 framework to be blocked in these situations. In addition, if the training data volume is insufficient or the data is too single, the model may also have overfitting problems. Therefore, to solve these problems in the electric vehicle recognition scenario when entering the elevator, it is necessary to improve on the basis of the traditional YOLOv8 framework to further enhance the robustness and efficiency of electric vehicle recognition and better cope with the challenges in this scenario. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides an elevator electric vehicle recognition method based on improved YOLOv8, aiming to solve the problems mentioned in the background art.
[0006] To achieve the above object, the present invention provides the following technical solution: An elevator electric vehicle recognition method based on improved YOLOv8, comprising the following steps:
[0007] Step S1: Construct an electric vehicle entering elevator image dataset, which includes several electric vehicle entering elevator images;
[0008] Step S2: Construct an electric vehicle recognition model, which consists of a feature extraction backbone network module, an aggregation neck network module, a multi-dimensional collaborative attention module, and a candidate box prediction module, and the feature extraction backbone network module, the aggregation neck network module, the multi-dimensional collaborative attention module, and the candidate box prediction module are arranged in series;
[0009] Step S3: Input the electric vehicle into-lift images in the electric vehicle into-lift image dataset into the feature extraction backbone network module for feature extraction, and respectively obtain a high-resolution electric vehicle into-lift feature map, a medium-resolution electric vehicle into-lift feature map, and a low-resolution electric vehicle into-lift feature map;
[0010] Step S4: Input the high-resolution electric vehicle into-lift feature map, the medium-resolution electric vehicle into-lift feature map, and the low-resolution electric vehicle into-lift feature map into the aggregation neck network module through parallel channels for fusion, and respectively obtain a high-resolution electric vehicle into-lift aggregation feature map, a medium-resolution electric vehicle into-lift aggregation feature map, and a low-resolution electric vehicle into-lift aggregation feature map;
[0011] Step S5: Input the high-resolution electric vehicle into-lift aggregation feature map, the medium-resolution electric vehicle into-lift aggregation feature map, and the low-resolution electric vehicle into-lift aggregation feature map into the multi-dimensional collaborative attention module. Extract dimension feature maps from the channels, space, and position dimensions of the high-resolution electric vehicle into-lift aggregation feature map, the medium-resolution electric vehicle into-lift aggregation feature map, and the low-resolution electric vehicle into-lift aggregation feature map respectively according to different resolutions. Pool the dimension feature maps through average pooling and standard deviation pooling, calculate the weights of the pooled dimension feature maps through excitation transformation, and generate a multi-dimensional weighted feature map;
[0012] Step S6: Input the multi-dimensional weighted feature map into the candidate box prediction module, perform object prediction on the multi-dimensional weighted feature map through bounding box regression, object classification, and confidence prediction, and post-process the prediction results to generate the final prediction output.
[0013] Further, the aggregation neck network module is composed of a first local aggregation module, a second local aggregation module, a third local aggregation module, a fourth local aggregation module, a first group exchange convolution block, a second group exchange convolution block, and a third group exchange convolution block;
[0014] The processing flow of the aggregated neck network module is as follows: Upsample the low-resolution electric vehicle entering the elevator feature map and splice it with the medium-resolution electric vehicle entering the elevator feature map to obtain a medium-low scale fusion feature map. Input the medium-low scale fusion feature map into the first local aggregation module to obtain a medium-low scale aggregated feature map. Input the medium-low scale aggregated feature map into the first group exchange convolution block to obtain a medium-low scale convolution feature map. Splice the medium-low scale convolution feature map with the high-resolution electric vehicle entering the elevator feature map to obtain a multi-scale aggregated feature map. Input the multi-scale aggregated feature map into the second local aggregation module to obtain a high-resolution electric vehicle entering the elevator aggregated feature map. Input the high-resolution electric vehicle entering the elevator aggregated feature map into the second group exchange convolution block to obtain a multi-scale convolution feature map. Splice the multi-scale convolution feature map with the medium-low scale aggregated feature map to obtain a multi-scale spliced feature map. Input the multi-scale spliced feature map into the third local aggregation module to obtain a medium-resolution electric vehicle entering the elevator aggregated feature map. Input the medium-resolution electric vehicle entering the elevator aggregated feature map into the third group exchange convolution block to obtain a multi-layer convolution feature map. Splice the multi-layer convolution feature map with the low-resolution electric vehicle entering the elevator feature map to obtain a multi-layer spliced feature map. Input the multi-layer spliced feature map into the fourth local aggregation module to obtain a low-resolution electric vehicle entering the elevator aggregated feature map;
[0015] Among them, the upsampling operation uses the bilinear interpolation method.
[0016] Furthermore, the feature extraction backbone network module consists of a first convolutional layer, a second convolutional layer, a first feature fusion and splicing layer, a third convolutional layer, a second feature fusion and splicing layer, a fourth convolutional layer, a third feature fusion and splicing layer, a fifth convolutional layer, a fourth feature fusion and splicing layer, and a pooling layer connected in sequence; among them, the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer have the same structure; the first feature fusion and splicing layer, the second feature fusion and splicing layer, and the third feature fusion and splicing layer have the same structure.
[0017] Furthermore, the processing flow of the feature extraction backbone network module is as follows: The input image, that is, the electric vehicle entering the elevator image, passes through the first convolutional layer, the second convolutional layer, the first feature fusion and splicing layer, the third convolutional layer, the second feature fusion and splicing layer, the fourth convolutional layer, the third feature fusion and splicing layer, the fifth convolutional layer, the fourth feature fusion and splicing layer, and the pooling layer in sequence. Among them, the output of the second feature fusion and splicing layer is the high-resolution electric vehicle entering the elevator feature map X b , the output of the third feature fusion and splicing layer is the medium-resolution electric vehicle entering the elevator feature map X m , and the output of the pooling layer is the low-resolution electric vehicle entering the elevator feature map X s .
[0018] Furthermore, the structures of the first local aggregation module, the second local aggregation module, the third local aggregation module, and the fourth local aggregation module are the same, and they are composed of the VoVNet module and the CSPNet module combined;
[0019] The processing flow of the first local aggregation module is as follows:
[0020] VoVNet module: First, perform multiple convolution operations on the input, that is, the medium and low-scale fusion feature map, to obtain the multi-convolution feature map X(X 1 conv , X 2 conv ), where X 1 conv is the feature map after the first ordinary convolution operation, and X 2 conv is the feature map obtained by performing the second ordinary convolution operation on X 1 conv ; Subsequently, perform a depthwise separable convolution operation on X 2 conv to obtain the depth convolution feature map X 3 conv ; Finally, perform a concatenation operation on the multi-convolution feature map X and the depth convolution feature map X 3 conv to obtain the output X 4 conv of the VoVNet module;
[0021] CSPNet module: First, use the splitting operation to divide the output X 4 conv of the VoVNet module into the first part X part1 and the second part X part2 , then perform a convolution operation on the second part X part2 to obtain the cross-stage convolution feature map X part2_conv , and then perform a concatenation operation on the first part X part1 and the cross-stage convolution feature map X part2_conv to obtain the output of the first local aggregation module.
[0022] Furthermore, the processing flow of the first convolutional layer is: Represent the input image, that is, the electric vehicle entering the elevator image, as , represents the number of channels of the input image, and respectively represent the height and width of the input image , represents the set of real numbers; The processing process of the first convolutional layer is represented as:
[0023] ;
[0024] In the formula, is the feature map output after the first convolutional layer; is the activation function SiLU; represents the convolutional operation; and are the weights and biases of the first convolutional layer; represents the step size of each slide during the convolutional operation.
[0025] Furthermore, the processing flow of the first feature fusion and splicing layer is as follows: After inputting the output of the second convolutional layer into the first feature fusion and splicing layer, it is divided into a direct transmission part and a convolutional processing part , which is expressed as:
[0026] ;
[0027] In the formula, represents the direct transmission part; represents the convolutional processing part; , and respectively represent and the number of channels of; represents the first feature fusion and splicing layer;
[0028] processes the convolutional processing part , which is expressed as:
[0029] ;
[0030] In the formula, , respectively represent the weights and biases of the th convolutional layer in the first feature fusion and splicing layer;
[0031] After processing is concatenated with to obtain the concatenated feature , and finally the concatenated feature passes through the final convolutional layer of the first feature fusion and splicing layer to obtain the output of the first feature fusion and splicing layer, which is expressed as:
[0032] ;
[0033] ;
[0034] In the formula, represents the concatenation operation; and Represent the weights and biases of the final convolutional layer of the first feature fusion and splicing layer.
[0035] Furthermore, the processing flow of the pooling layer is as follows: Input the output of the fourth feature fusion and splicing layer into the pooling layer, first pass through a convolutional layer with a stride of 1 to obtain the activation feature map , and then perform a max pooling operation on the activation feature map to obtain the pooled feature map
[0036] ;
[0037] In the formula, represents the feature map generated from pooling kernels k of different sizes in the pooling layer, that is, the pooled feature map; represents performing a multi-scale max pooling operation on the activation feature map ; represents the size of the pooling kernel; represents the padding number for the activation feature map ; is the specific value of the padding number;
[0038] Concatenate the activation feature map and the pooled feature map on the channel dimension, and integrate the multi-scale information through a convolutional operation, and at the same time compress the number of channels to obtain the output of the pooling layer, that is, the low-resolution feature map X s .
[0039] Furthermore, the structures of the first group shuffle convolutional block, the second group shuffle convolutional block, and the third group shuffle convolutional block are the same;
[0040] The working process of the first group shuffle convolutional block is as follows: First, perform a standard convolutional operation on the input, that is, the output of the first local aggregation module, through a convolutional layer to obtain the group convolutional feature map X csp1 , then, pass the group convolution X csp1 through a 1×1 convolutional layer for another convolutional operation to generate additional ghost features and obtain the ghost feature map X ghost , then, shuffle the channel order of the ghost features through a channel shuffle operation to obtain the group convolutional ghost feature map , that is, the output of the first group shuffle convolutional block, which is represented as:
[0041] ;
[0042] In the formula, X ghost_s is the group convolutional ghost feature map obtained after the channel shuffle operation on the ghost features; represents the channel shuffle operation.
[0043] Further, the specific process of step S5 is as follows: The output of the aggregated neck network module , that is, the high-resolution electric vehicle entering ladder aggregated feature map, the medium-resolution electric vehicle entering ladder aggregated feature map, and the low-resolution electric vehicle entering ladder aggregated feature map, are input into the multi-dimensional collaborative attention module, and the attention weights are calculated and collaboratively fused through channel dimension attention feature modeling, spatial dimension attention feature modeling, and position dimension attention feature modeling respectively;
[0044] The specific process of channel dimension attention feature modeling is as follows: First, for the input, that is, the output of the aggregated neck network module perform global average pooling in the spatial dimension to obtain the channel weight vector f c ∈R c , expressed as:
[0045] ;
[0046] In the formula, represents the average of all pixel values on channel c; respectively represent the indices of rows and columns, and the value ranges are and ;
[0047] Through double summation traverse all spatial positions, that is, all pixel points, on channel c, accumulate the pixel values, and obtain X p (c, u, v), X p (c, u, v) represents the pixel value of the input at position (u, v) on channel c;
[0048] Perform a linear transformation on the channel weight vector f c to obtain the weight W c for each channel, and then apply W c to the output of each channel of to obtain the output of channel dimension attention feature modeling , expressed as:
[0049] ;
[0050] In the formula, is the activation function Sigmoid; W c and b c respectively represent the weight parameter and bias parameter for each channel;
[0051] The specific process of spatial dimension attention feature modeling is: For the input Aggregate using average pooling and max pooling in the channel dimension to obtain a two-dimensional spatial feature map , expressed as:
[0052] ;
[0053] In the formula, represents average pooling of the input in the channel dimension; represents max pooling of the input in the channel dimension; represents the pixel values of all channels at position (u, v) of the input ; represents all channels;
[0054] Multiply the two-dimensional spatial feature map by the input after learning the spatial weights through a 7×7 convolutional layer to obtain the output of spatial dimension attention feature modeling, expressed as:
[0055] ;
[0056] In the formula, represents performing a convolution operation on the two-dimensional spatial feature map using a 7×7 convolutional layer;
[0057] The specific process of position dimension attention feature modeling is as follows: Map the input to three different spatial generated vectors, expressed as:
[0058] ;
[0059] In the formula, represents the query vector; is the weight matrix of the query vector; represents the key vector; represents the value vector; is the weight matrix of the key vector; is the weight matrix of the value vector;
[0060] After obtaining the vectors, calculate the attention weights through the softmax function and perform weighting using the attention weights to obtain the output of position dimension attention feature modeling, expressed as:
[0061] ;
[0062] In the formula, represents the position in the input The query vector; Indicates the input at the position of the key vector; Indicates the input at the position of the value vector at that position; Indicates the Softmax function; Indicates the sum over all positions in the input; respectively indicate the input where the abscissa is and the ordinate is at the position;
[0063] Fuse , , to form the final multi-dimensional weighted feature map , expressed as:
[0064] ;
[0065] In the formula, represents element-wise multiplication.
[0066] Compared with the existing technology, the present invention has the following beneficial effects:
[0067] (1) By designing a local aggregation module and introducing it into the aggregation neck network module of the electric vehicle recognition model, the electric vehicle recognition model can perform better in complex scenarios through hierarchical feature aggregation, and the recognition ability for small electric vehicles in the elevator and the edge details of electric vehicles is significantly improved; by combining convolutional linear transformation and channel shuffle operations and introducing them into the electric vehicle recognition model together with the local aggregation module, the recognition accuracy of the electric vehicle recognition model is greatly improved.
[0068] (2) By introducing a multi-dimensional collaborative attention module into the electric vehicle recognition model, the electric vehicle recognition model can consider feature interactions at different levels and can simultaneously process multiple input information. By combining this information, the electric vehicle recognition model has the ability of efficient spatial-scale joint processing, can more accurately locate the target during the detection process, and can also better recognize and distinguish electric vehicles from other objects in complex environments such as light changes, partial occlusion, or complex backgrounds. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a schematic structural diagram of the electric vehicle recognition model of the present invention.
[0070] Figure 2Schematic diagram of the feature extraction backbone network module of the present invention. Detailed implementation manners
[0071] As Figure 1 shown, the present invention provides a technical solution: an elevator-internal electric vehicle recognition method based on improved YOLOv8, including the following steps:
[0072] Step S1: Construct an electric vehicle entering elevator image dataset, which includes a number of electric vehicle entering elevator images.
[0073] The specific process of step S1 is as follows: Since there is a lack of a publicly available elevator-internal electric vehicle image dataset, more than 6,000 images of electric vehicles entering elevators (electric vehicle entering elevator images) are manually collected. These images are from the cameras of various surrounding communities and residential buildings as well as artificial shooting; the images cover various types of electric vehicles, including traditional electric vehicles and new types of electric vehicles for the elderly, etc.
[0074] Preprocess the collected electric vehicle entering elevator images:
[0075] 1. Remove similar elevator-internal electric vehicle images.
[0076] 2. Semi-automatically annotate the electric vehicles in the collected electric vehicle entering elevator images using an annotation tool (label-studio).
[0077] 3. Conduct a consistency check on the annotated electric vehicle entering elevator images, that is, calculate the similarity of the annotation results of multiple people for the same content, and take the annotation result with a high similarity to ensure the rationality of the annotation.
[0078] Construct an electric vehicle entering elevator image dataset from the preprocessed electric vehicle entering elevator images.
[0079] Step S2: Construct an electric vehicle recognition model, which is composed of a feature extraction backbone network module, an aggregation neck network module, a multi-dimensional collaborative attention module, and a candidate box prediction module, and the feature extraction backbone network module, the aggregation neck network module, the multi-dimensional collaborative attention module, and the candidate box prediction module are arranged in series.
[0080] Step S3: Input the electric vehicle entering elevator images in the electric vehicle entering elevator image dataset into the feature extraction backbone network module for feature extraction, and respectively obtain a high-resolution electric vehicle entering elevator feature map, a medium-resolution electric vehicle entering elevator feature map, and a low-resolution electric vehicle entering elevator feature map.
[0081] As Figure 2As shown in the figure, the feature extraction backbone network module consists of a first convolutional layer (CBS), a second convolutional layer, a first feature fusion and splicing layer (C2f), a third convolutional layer, a second feature fusion and splicing layer, a fourth convolutional layer, a third feature fusion and splicing layer, a fifth convolutional layer, a fourth feature fusion and splicing layer, and a pooling layer (SPPF) connected in sequence; among them, the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer, and the fifth convolutional layer have the same structure; the first feature fusion and splicing layer, the second feature fusion and splicing layer, and the third feature fusion and splicing layer have the same structure.
[0082] Among them, the processing flow of the feature extraction backbone network module is as follows: the input image (electric vehicle entering the elevator image) passes through the first convolutional layer, the second convolutional layer, the first feature fusion and splicing layer, the third convolutional layer, the second feature fusion and splicing layer, the fourth convolutional layer, the third feature fusion and splicing layer, the fifth convolutional layer, the fourth feature fusion and splicing layer, and the pooling layer in sequence. Among them, the output of the second feature fusion and splicing layer is the high-resolution electric vehicle entering the elevator feature map X b , the output of the third feature fusion and splicing layer is the medium-resolution electric vehicle entering the elevator feature map X m , and the output of the pooling layer is the low-resolution electric vehicle entering the elevator feature map X s .
[0083] Among them, the processing flow of the first convolutional layer is as follows: the input image (electric vehicle entering the elevator image) is represented as , represents the number of channels of the input image, and respectively represent the height and width of the input image , represents the set of real numbers; the input image first effectively extracts image features through the first convolutional layer, reducing the spatial dimension of the image for better and faster image processing; the processing process of the first convolutional layer can be expressed as:
[0084] (1);
[0085] In the formula, is the feature map output after passing through the first convolutional layer; is the activation function SiLU; represents the convolution operation; and are the weights and biases of the first convolutional layer; represents the step size of each slide during the convolution operation.
[0086] Among them, the processing flow of the first feature fusion and splicing layer is as follows: after the output of the second convolutional layer is input into the first feature fusion and splicing layer, it is divided into a direct transmission part And the convolution processing part , which is expressed as:
[0087] (2);
[0088] In the formula, represents the direct transmission part without convolution operation; represents the convolution processing part for subsequent convolution processing; , and respectively represent and the number of channels of; represents the first feature fusion and splicing layer.
[0089] Process the convolution processing part , which is expressed as:
[0090] (3);
[0091] In the formula, , respectively represent the weights and biases of the th convolution layer in the first feature fusion and splicing layer.
[0092] After processing the and are spliced to obtain the spliced feature , and finally the spliced feature passes through the final convolution layer of the first feature fusion and splicing layer to obtain the output of the first feature fusion and splicing layer, which is expressed as:
[0093] (4);
[0094] (5);
[0095] In the formula, represents the splicing operation; and represent the weights and biases of the final convolution layer of the first feature fusion and splicing layer.
[0096] The pooling layer (SPPF) can significantly expand the receptive field by combining convolution operations and max-pooling operations, helping the electric vehicle recognition model to more accurately recognize small target groups.
[0097] The processing flow of the pooling layer (SPPF) is as follows: The output of the fourth feature fusion and splicing layer is input into the pooling layer (SPPF), and first passes through a convolution layer with a stride of 1 (used to extract preliminary features and adjust the number of channels) to obtain the activation feature map , and then perform max pooling operation on the activated feature map to further increase the receptive field and obtain the pooled feature map , which is expressed as:
[0098] (6);
[0099] In the formula, represents the feature map generated by pooling kernels k of different sizes in the pooling layer, that is, the pooled feature map; represents performing multi-scale max pooling operation on the activated feature map ; represents the size of the pooling kernel; represents the padding number for the activated feature map ; is the specific value of the padding number.
[0100] Concatenate the activated feature map and the pooled feature map along the channel dimension, and integrate multi-scale information through convolution operation, and at the same time compress the number of channels to obtain the output of the pooling layer, that is, the low-resolution feature map X s .
[0101] Step S4: Input the high-resolution electric vehicle into-lift feature map, medium-resolution electric vehicle into-lift feature map, and low-resolution electric vehicle into-lift feature map into the aggregation neck network module through parallel channels for fusion, and respectively obtain the high-resolution electric vehicle into-lift aggregation feature map, medium-resolution electric vehicle into-lift aggregation feature map, and low-resolution electric vehicle into-lift aggregation feature map.
[0102] Among them, the aggregation neck network module is composed of a first local aggregation module, a second local aggregation module, a third local aggregation module, a fourth local aggregation module, a first group exchange convolution block, a second group exchange convolution block, and a third group exchange convolution block.
[0103] The processing flow of the aggregated neck network module is as follows: Upsample the low-resolution electric vehicle entering the ladder feature map and splice it with the medium-resolution electric vehicle entering the ladder feature map to obtain a medium-low scale fusion feature map. Input the medium-low scale fusion feature map into the first local aggregation module to obtain a medium-low scale aggregated feature map. Input the medium-low scale aggregated feature map into the first group exchange convolution block to obtain a medium-low scale convolution feature map. Splice the medium-low scale convolution feature map with the high-resolution electric vehicle entering the ladder feature map to obtain a multi-scale aggregated feature map. Input the multi-scale aggregated feature map into the second local aggregation module to obtain a high-resolution electric vehicle entering the ladder aggregated feature map. Input the high-resolution electric vehicle entering the ladder aggregated feature map into the second group exchange convolution block to obtain a multi-scale convolution feature map. Splice the multi-scale convolution feature map with the medium-low scale aggregated feature map to obtain a multi-scale spliced feature map. Input the multi-scale spliced feature map into the third local aggregation module to obtain a medium-resolution electric vehicle entering the ladder aggregated feature map. Input the medium-resolution electric vehicle entering the ladder aggregated feature map into the third group exchange convolution block to obtain a multi-layer convolution feature map. Splice the multi-layer convolution feature map with the low-resolution electric vehicle entering the ladder feature map to obtain a multi-layer spliced feature map. Input the multi-layer spliced feature map into the fourth local aggregation module to obtain a low-resolution electric vehicle entering the ladder aggregated feature map.
[0104] Among them, the upsampling operation adopts the bilinear interpolation method, which estimates new pixel values by linearly weighting the surrounding pixels, so as to achieve a smoother spatial change response.
[0105] Among them, the structures of the first local aggregation module (VoVGSCSPG), the second local aggregation module, the third local aggregation module, and the fourth local aggregation module are all the same, and they are composed of the VoVNet module and the CSPNet module.
[0106] The processing flow of the first local aggregation module is as follows:
[0107] The VoVNet module is used for progressive feature extraction; in order to hierarchically extract the input features, the present invention utilizes the progressive connection structure of the VoVNet module; first, perform multiple convolution operations on the input (medium-low scale fusion feature map) to obtain a multi-convolution feature map X(X 1 conv ,X 2 conv ), where X 1 conv is the feature map after the first ordinary convolution operation, and X 2 conv is the feature map after performing the second ordinary convolution operation on X 1 conv ; subsequently, perform a depthwise separable convolution operation on X 2 conv to obtain a depth convolution feature map X3 conv ; Finally, concatenate the multi-convolution feature map X and the depth convolution feature map X 3 conv to obtain the output X of the VoVNet module 4 conv .
[0108] The CSPNet module is used for cross-stage partial fusion; to reduce feature redundancy and optimize gradient flow, the CSPNet module improves computational efficiency through feature map segmentation and partial fusion. First, use the segmentation operation to divide the output X of the VoVNet module 4 conv into a first part X part1 and a second part X part2 . Then, perform a convolution operation on the second part X part2 to obtain the cross-stage convolution feature map X part2_conv . Next, concatenate the first part X part1 and the cross-stage convolution feature map X part2_conv to obtain the output of the first local aggregation module.
[0109] Among them, the first group ghost convolution block (GHConv), the second group ghost convolution block, and the third group ghost convolution block have the same structure.
[0110] The working process of the first group ghost convolution block is as follows: to reduce the amount of computation and improve the feature fusion effect, promote the information flow between features by shuffling the channel order. First, perform a standard convolution operation on the input (the output of the first local aggregation module) through a convolutional layer to extract feature information to obtain the group convolution feature map X csp1 . Then, perform a convolution operation on the group convolution X csp1 through a 1×1 convolutional layer to generate additional ghost features and obtain the ghost feature map X ghost . Then, shuffle the channel order of the ghost features through the channel shuffle operation to improve the feature expression ability of the network and obtain the group convolution ghost feature map , that is, the output of the first group ghost convolution block, denoted as:
[0111] (7);
[0112] In the formula, X ghost_s is the group convolution ghost feature map obtained after the channel shuffle operation on the ghost features, and the size remains unchanged; represents the channel shuffle operation.
[0113] Step S5: Input the high-resolution electric vehicle into-ladder aggregation feature map, the medium-resolution electric vehicle into-ladder aggregation feature map, and the low-resolution electric vehicle into-ladder aggregation feature map into the multi-dimensional collaborative attention module. Dimension feature maps are extracted from the high-resolution electric vehicle into-ladder aggregation feature map, the medium-resolution electric vehicle into-ladder aggregation feature map, and the low-resolution electric vehicle into-ladder aggregation feature map respectively according to different resolutions in three dimensions: channels, space, and position. The dimension feature maps are aggregated through average pooling and standard deviation pooling, and the weights of the aggregated dimension feature maps are calculated through excitation transformation to generate multi-dimensional weighted feature maps.
[0114] Among them, the multi-dimensional collaborative attention module focuses on feature modeling in three major dimensions: channel dimension attention feature modeling, spatial dimension attention feature modeling, and position dimension attention feature modeling. By calculating attention weights separately in these three dimensions and performing collaborative fusion, this module can effectively strengthen the key information in the input feature map.
[0115] The specific process of channel dimension attention feature modeling is as follows: First, perform global average pooling on the input, that is, the output of the aggregation neck network module (the high-resolution electric vehicle into-ladder aggregation feature map, the medium-resolution electric vehicle into-ladder aggregation feature map, and the low-resolution electric vehicle into-ladder aggregation feature map) in the spatial dimension to obtain the channel weight vector f c ∈R c , which is expressed as:
[0116] (8);
[0117] In the formula, represents the average of all pixel values on channel c; respectively represent the indices of the row (height) and column (width) in and .
[0118] Through double summation traverse all spatial positions (i.e., all pixel points) on channel c, accumulate the pixel values, and obtain X p (c, u, v), X p (c, u, v) represents the pixel value of the input at channel c and position (u, v).
[0119] After obtaining the channel weight vector f c , then obtain the weight W c of each channel through a simple linear transformation, and then apply W c to the output of each channel of to obtain the output of channel dimension attention feature modeling, which is expressed as:
[0120] (9);
[0121] Wherein, is the activation function Sigmoid; W c and b c represent the weight parameter and bias parameter of each channel respectively.
[0122] The specific process of spatial dimension attention feature modeling is as follows: Aggregate the input in the channel dimension using average pooling and max pooling to obtain a two-dimensional spatial feature map , expressed as:
[0123] (10);
[0124] Wherein, represents average pooling of the input in the channel dimension; represents max pooling of the input in the channel dimension; represents the pixel value of all channels of the input at the position (u, v); represents all channels.
[0125] Subsequently, multiply the two-dimensional spatial feature map by the input after learning the spatial weight through a 7×7 convolutional layer to obtain the output of the spatial dimension attention feature modeling, expressed as:
[0126] (11);
[0127] Wherein, represents performing a convolution operation on the two-dimensional spatial feature map using a 7×7 convolutional layer.
[0128] The specific process of position dimension attention feature modeling is as follows: Map the input to three different spaces to generate vectors, expressed as:
[0129] (12);
[0130] Wherein, represents the query vector; is the weight matrix of the query vector; represents the key vector; represents the value vector; is the weight matrix of the key vector; is the weight matrix of the value vectors.
[0131] After obtaining the vectors, the attention weights are calculated through the softmax function, and these attention weights are used for weighting to obtain the output of the position dimension attention feature modeling , expressed as:
[0132] (13);
[0133] In the formula, represents the input at position of the query vector, which is used to find other positions in the input related to position ; represents the key vector at position in the input , which is used to describe the features of each position in the input ; represents the value vector at position in the input ; represents the Softmax function; represents the sum over all positions in the input ; respectively represent the position in the input with abscissa and ordinate .
[0134] Fuse , , to form the final multi-dimensional weighted feature map , expressed as:
[0135] (14);
[0136] In the formula, represents element-wise multiplication.
[0137] Step S6: Input the multi-dimensional weighted feature map into the candidate box prediction module, perform object prediction on the multi-dimensional weighted feature map through bounding box regression, object classification, and confidence prediction, and post-process the prediction results to generate the final prediction output.
[0138] Among them, the candidate box prediction module uses the detection head of the original YOLOv8.
[0139] In the loss function of bounding box regression, the traditional CIoU loss function only focuses on the overlapping area between the predicted box and the ground truth box, and has insufficient optimization for the aspect ratio, making it unable to effectively constrain the shape and size of the predicted box. Therefore, the present invention adopts the MPDIoU loss function as the loss function for bounding box regression to achieve a more comprehensive position metric and more effective gradient information; the MPDIoU loss function is expressed as:
[0140] (15);
[0141] In the formula, represents the intersection over union of the predicted box and the ground truth box; and are additional weight coefficients used to balance the influence of the new distance metric and the aspect ratio; represents the maximum distance between the predicted box and the ground truth box considering the influence of the direction angle in polar coordinates; represents the maximum value of the diagonal lengths of the predicted box and the ground truth box; is a term measuring the difference in aspect ratio between the predicted box and the ground truth box, is the weight coefficient used to balance the aspect ratio in the original CIoU loss function, and here, through the weight coefficient and the term measuring the difference in aspect ratio between the predicted box and the ground truth box are multiplied to achieve a more comprehensive position metric.
[0142] Among them, the electric vehicle recognition model uses the processed electric vehicle entering the elevator image dataset to train the object detection model in a supervised learning manner. All parameters are learnable parameters, and are optimized through the backpropagation algorithm in combination with the MPDIoU loss function, cross-entropy loss function, and binary cross-entropy loss function; during the training process, the number of iterations of the electric vehicle recognition model is set to 100, the learning rate is 0.01, and the SGD optimizer is used to optimize the parameters of the electric vehicle recognition model; after all training is completed, the parameters of the electric vehicle recognition model with the best performance are saved, and finally, electric vehicle recognition prediction and performance evaluation are carried out on the test set; the test set is obtained by dividing the electric vehicle entering the elevator image dataset.
[0143] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. The electric vehicle identification method in the elevator based on the improved YOLOv8 is characterized by: The steps include: Step S1: constructing an electric vehicle entering an elevator image dataset, wherein the electric vehicle entering an elevator image dataset includes a number of electric vehicle entering an elevator images; Step S2: constructing an electric vehicle recognition model, wherein the electric vehicle recognition model is composed of a feature extraction backbone network module, an aggregation neck network module, a multi-dimensional collaborative attention module and a candidate box prediction module, and the feature extraction backbone network module, the aggregation neck network module, the multi-dimensional collaborative attention module and the candidate box prediction module are arranged in series; Step S3: inputting the electric vehicle entering the elevator image in the electric vehicle entering the elevator image data set into the feature extraction backbone network module for feature extraction, and obtaining a high-resolution electric vehicle entering the elevator feature map, a medium-resolution electric vehicle entering the elevator feature map, and a low-resolution electric vehicle entering the elevator feature map respectively; Step S4: inputting the high-resolution electric vehicle entering the elevator feature map, the medium-resolution electric vehicle entering the elevator feature map and the low-resolution electric vehicle entering the elevator feature map into the aggregation neck network module through parallel channels for fusion, thereby obtaining a high-resolution electric vehicle entering the elevator aggregation feature map, a medium-resolution electric vehicle entering the elevator aggregation feature map and a low-resolution electric vehicle entering the elevator aggregation feature map respectively; Step S5: inputting the high-resolution electric vehicle entry aggregated feature map, the medium-resolution electric vehicle entry aggregated feature map and the low-resolution electric vehicle entry aggregated feature map into the multidimensional collaborative attention module, extracting dimensional feature maps in three dimensions of channel, space and position from the high-resolution electric vehicle entry aggregated feature map, the medium-resolution electric vehicle entry aggregated feature map and the low-resolution electric vehicle entry aggregated feature map according to different resolutions, aggregating the dimensional feature maps through average pooling and standard deviation pooling, calculating the weights of the aggregated dimensional feature maps through excitation transformation, and generating a multidimensional weighted feature map; Step S6: input the multidimensional weighted feature map into the candidate box prediction module, perform target prediction on the multidimensional weighted feature map through bounding box regression, target classification and confidence prediction, post-process the prediction results, and generate the final prediction output; The aggregation neck network module is composed of a first local aggregation module, a second local aggregation module, a third local aggregation module, a fourth local aggregation module, a first group exchange convolution block, a second group exchange convolution block and a third group exchange convolution block; The processing flow of the aggregation neck network module is as follows: upsampling the low-resolution electric vehicle entry feature map and splicing it with the medium-resolution electric vehicle entry feature map to obtain a medium- and low-scale fusion feature map, inputting the medium- and low-scale fusion feature map into the first local aggregation module to obtain a medium- and low-scale aggregation feature map, inputting the medium- and low-scale aggregation feature map into the first group exchange convolution block to obtain a medium- and low-scale convolution feature map, splicing the medium- and low-scale convolution feature map with the high-resolution electric vehicle entry feature map to obtain a multi-scale aggregation feature map, inputting the multi-scale aggregation feature map into the second local aggregation module to obtain a high-resolution electric vehicle entry aggregation feature map, and inputting the high-resolution electric vehicle entry feature map into the second local aggregation module. The electric vehicle entering the elevator aggregate feature map is input into the second group exchange convolution block to obtain a multi-scale convolution feature map, the multi-scale convolution feature map is spliced with the medium and low-scale aggregate feature map to obtain a multi-scale spliced feature map, the multi-scale spliced feature map is input into the third local aggregation module to obtain a medium-resolution electric vehicle entering the elevator aggregate feature map, the medium-resolution electric vehicle entering the elevator aggregate feature map is input into the third group exchange convolution block to obtain a multi-layer convolution feature map, the multi-layer convolution feature map is spliced with the low-resolution electric vehicle entering the elevator feature map to obtain a multi-layer spliced feature map, the multi-layer spliced feature map is input into the fourth local aggregation module to obtain a low-resolution electric vehicle entering the elevator aggregate feature map; Among them, the upsampling operation adopts the bilinear interpolation method; The feature extraction backbone network module is composed of a first convolutional layer, a second convolutional layer, a first feature fusion splicing layer, a third convolutional layer, a second feature fusion splicing layer, a fourth convolutional layer, a third feature fusion splicing layer, a fifth convolutional layer, a fourth feature fusion splicing layer and a pooling layer, which are connected in sequence; wherein, the first convolutional layer, the second convolutional layer, the third convolutional layer, the fourth convolutional layer and the fifth convolutional layer have the same structure; the first feature fusion splicing layer, the second feature fusion splicing layer and the third feature fusion splicing layer have the same structure; The processing flow of the feature extraction backbone network module is as follows: the input image, i.e., the electric vehicle entering the elevator image, passes through the first convolution layer, the second convolution layer, the first feature fusion splicing layer, the third convolution layer, the second feature fusion splicing layer, the fourth convolution layer, the third feature fusion splicing layer, the fifth convolution layer, the fourth feature fusion splicing layer and the pooling layer in sequence, wherein the output of the second feature fusion splicing layer is the high-resolution electric vehicle entering the elevator feature map X b The output of the third feature fusion splicing layer is the medium-resolution electric vehicle entry feature map X m The output of the pooling layer is the low-resolution electric vehicle entry feature map X s ; The first local aggregation module, the second local aggregation module, the third local aggregation module and the fourth local aggregation module have the same structure, and are formed by combining the VoVNet module and the CSPNet module; The processing flow of the first local aggregation module is: VoVNet module: First, multiple convolution operations are performed on the input, i.e., the medium and low scale fusion feature maps, to obtain a multi-convolution feature map X(X 1 conv ,X 2 conv ), where X 1 conv is the feature map after the first ordinary convolution operation, X 2 conv Is X 1 conv The feature map after the second ordinary convolution operation; then, X 2 conv After the depth-separable convolution operation, the depth convolution feature map X is obtained 3 conv ; Finally, the multi-convolution feature map X and the deep convolution feature map X 3 conv Perform concatenation to get the output X of the VoVNet module 4 conv ; CSPNet module: First, the output X of the VoVNet module is split using a split operation. 4 conv Divided into Part I X part1 and Part II X part2 , then, for the second part X part2 Perform convolution operation to obtain the cross-stage convolution feature map X part2_conv , and then the first part X part1 and the cross-stage convolutional feature map X part2_conv Perform a splicing operation to obtain the output of the first local aggregation module; The processing flow of the first convolutional layer is: the input image, i.e. the image of the electric car entering the elevator, is represented as , Indicates the number of channels of the input image, and Represents the input image The height and width of represents a set of real numbers; the processing process of the first convolutional layer is expressed as: ; In the formula, It is the feature map output after the first convolutional layer; is the activation function SiLU; represents the convolution operation; and are the weights and biases of the first convolutional layer; Indicates the step size of each slide during the convolution operation; The processing flow of the first feature fusion splicing layer is: the output of the second convolutional layer After inputting the first feature fusion concatenation layer, it is divided into the direct transfer part And the convolution processing part , expressed as: ; In the formula, Indicates the direct delivery part; Represents the convolution processing part; , and Respectively and The number of channels; represents the first feature fusion splicing layer; Convolution processing part Processing is performed, expressed as: ; In the formula, , They represent the first feature fusion concatenation layer. The weights and biases of the convolutional layers; After processing and Perform splicing to obtain splicing features , and finally the splicing features After passing through the final convolution layer of the first feature fusion splicing layer, the output of the first feature fusion splicing layer is obtained , expressed as: ; ; In the formula, Represents a splicing operation; and Represents the weights and biases of the final convolutional layer of the first feature fusion concatenation layer; The processing flow of the pooling layer is as follows: the output of the fourth feature fusion splicing layer is input into the pooling layer, and first passes through a convolutional layer with a step size of 1 to obtain the activation feature map , and then activate the feature map Perform the maximum pooling operation to obtain the pooling feature map , expressed as: ; In the formula, Represents the feature map generated from pooling kernels k of different sizes in the pooling layer, that is, the pooling feature map; Represents the activation feature map Perform multi-scale maximum pooling operations; Indicates the size of the pooling kernel; Represents the activation feature map The number of fills; is the specific value of the filling quantity; The feature map will be activated With pooling feature map The concatenation is performed in the channel dimension, and the multi-scale information is integrated through the convolution operation, and the number of channels is compressed at the same time to obtain the output of the pooling layer, that is, the low-resolution feature map X s ; The first group exchange convolution block, the second group exchange convolution block, and the third group exchange convolution block have the same structure; The workflow of the first group exchange convolution block is as follows: First, the input, i.e., the output of the first local aggregation module, is first subjected to a standard convolution operation through a convolution layer to obtain the group convolution feature map X csp1 , then, the group convolution X csp1 Then perform a convolution operation through a 1×1 convolution layer to generate additional ghost features and obtain the ghost feature map X ghost Then, the channel order of the ghost feature is disrupted by the channel shuffle operation to obtain the group convolution ghost feature map , which is the output of the first group exchange convolution block, is expressed as: ; In the formula, X ghost_s It is the group convolution ghost feature map obtained after the ghost feature is shuffled through the channels; Represents a channel shuffle operation; The specific process of step S5 is: aggregate the output of the neck network module That is, the high-resolution electric vehicle entry aggregate feature map, the medium-resolution electric vehicle entry aggregate feature map and the low-resolution electric vehicle entry aggregate feature map are input into the multi-dimensional collaborative attention module, and the attention weights are calculated and collaboratively fused through channel dimension attention feature modeling, space dimension attention feature modeling and position dimension attention feature modeling respectively; The specific process of channel dimension attention feature modeling is as follows: First, the input is the output of the aggregation neck network module Perform global average pooling in the spatial dimension to obtain the channel weight vector f c ∈R c , expressed as: ; In the formula, Represents the average of all pixel values on channel c; Respectively The row and column indexes in the range and ; By double summation Traverse all spatial positions on channel c, that is, all pixel points, accumulate the pixel values, and get X p (c,u,v),X p (c,u,v) represents input The pixel value at position (u,v) on channel c; For the channel weight vector f c Perform a linear transformation to obtain the weight W of each channel c , then W c Application Each channel output of , get the output of channel dimension attention feature modeling , expressed as: ; In the formula, is the activation function Sigmoid; W c and b c Represent the weight parameters and bias parameters of each channel respectively; The specific process of modeling spatial dimension attention features is as follows: Aggregate using average pooling and maximum pooling in the channel dimension to obtain a two-dimensional spatial feature map , expressed as: ; In the formula, Represents the input in the channel dimension Perform average pooling; Represents the input in the channel dimension Perform maximum pooling; Represents input The pixel value on all channels at position (u,v); Indicates all channels; The two-dimensional spatial feature map After learning the spatial weights through a 7×7 convolutional layer, the input Multiply them together to get the output of spatial dimension attention feature modeling , expressed as: ; In the formula, Indicates the use of a 7×7 convolutional layer for the two-dimensional spatial feature map Perform convolution operation; The specific process of modeling the position dimension attention feature is as follows: Mapping to three different spaces generates vectors, expressed as: ; In the formula, represents the query vector; is the weight matrix of the query vector; represents the key vector; represents a value vector; is the weight matrix of the key vector; is the weight matrix of the value vector; After obtaining the vector, the attention weight is calculated through the softmax function, and the attention weight is used for weighting to obtain the output of the position dimension attention feature modeling , expressed as: ; In the formula, Represents input Middle position The query vector of Represents input Middle position The key vector of ; Represents input Middle position The value vector at ; Represents the Softmax function; Indicates input All locations Perform summation; Respectively represent input The horizontal coordinate is , the vertical axis is location; Will , , Fusion to form the final multi-dimensional weighted feature map , expressed as: ; In the formula, Represents element-wise multiplication; The MPDIoU loss function is used as the loss function for bounding box regression. , expressed as: ; In the formula, Represents the intersection-over-union ratio of the predicted box and the true box; and is an additional weight coefficient used to balance the impact of the new distance metric and aspect ratio; In polar coordinates, considering the direction angle The influence of, the maximum distance between the predicted box and the true box; Indicates the maximum diagonal length of the predicted box and the true box; It is a term that measures the difference in aspect ratio between the predicted box and the real box. It is the weight coefficient used to balance the aspect ratio in the original CIoU loss function.
Citation Information
Patent Citations
Medical image depth segmentation method based on fuzzy logic
CN116188435A
Fan surface defect detection method, device and equipment based on lightweight PC-EMA algorithm and storage medium
CN118691573A