Information fusion parking space detection method and system based on deep learning
Through the information fusion parking space detection method based on deep learning, the problems of low accuracy and poor efficiency of traditional detection methods are solved, and high-precision and high-efficiency parking space detection is achieved, which is suitable for modern urban intelligent transportation systems.
Patent Information
- Application Number
- CN202510009786.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional parking space detection methods have problems such as low accuracy, poor efficiency and poor adaptability, which are difficult to meet the needs of modern urban intelligent transportation systems.
The information fusion parking space detection method based on deep learning is adopted. By dividing the surround-view monitoring image into grid units, a feature map is extracted using a pre-trained convolutional neural network, and combining global and local information to integrate non-maximum suppression technology to generate the final parking space detection result.
It improves the accuracy and efficiency of parking space inspection, has significant advantages such as high detection rate, high accuracy and rapid detection, and has excellent performance under various environmental conditions.
Smart Images

Figure CN119942496A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the technical field of parking space detection, and in particular, to an information fusion parking space detection method and system based on deep learning. Background Art
[0002] With the acceleration of urbanization, the number of cars has increased year by year, and parking problems have become a common problem faced by many cities. The contradiction between parking supply and demand has become increasingly prominent. How to improve the management efficiency and detection accuracy of parking spaces has become an important issue that needs to be solved in the field of traffic management.
[0003] Traditional parking space detection methods mostly rely on manual inspections, geomagnetic induction or infrared sensors and other technologies. However, these methods often have a series of problems such as low accuracy, poor efficiency, and poor adaptability, and are difficult to meet the needs of modern urban intelligent transportation systems.
[0004] In recent years, although deep learning technology has made significant progress in the field of object detection, its application in parking space detection still needs to be further improved. Summary of the invention
[0005] In view of this, an embodiment of the present invention provides an information fusion parking space detection method and system based on deep learning, which aims to improve the accuracy and efficiency of parking space detection and has unique innovations and practical application value.
[0006] According to a first aspect of an embodiment of the present invention, there is provided an information fusion parking space detection method based on deep learning, comprising: step S1, dividing a collected 416×416 pixel surround monitoring image into a plurality of 13×13 grid units, and setting an 8-pixel overlapping area between adjacent grid units; step S2, performing multi-layer convolution and pooling processing on each grid unit through a pre-trained convolutional neural network to extract a feature map of size 13×13×512; step S3, applying a specified number of 3×3×512 convolution filters to the feature map, and extracting global information of the parking space through a Sigmoid function calculation, wherein the parking space entrance position is calculated by comparing the pixel mean of a specified area in the feature map with a preset threshold value. The type is determined by Softmax classification based on the pixel value distribution of different channels in the feature map, the occupancy status is judged according to the relationship between the pixel brightness of the specified area in the feature map and the set brightness threshold, and the local information of the parking space is extracted by a calculation method based on geometric relationships, including: determining the position of the connection point by calculating the coordinates of the specified key point in the feature map, and determining the direction by calculating the vector direction of adjacent key points; step S4, when the distance between the connection point in the global information and the connection point in the local information is less than 20 pixels, the non-maximum suppression technology based on the connection point is used to replace the global connection point with the local connection point, and the global information and local information of the parking space are integrated to generate the final parking space detection result.
[0007] In one implementation, the surround monitoring image of step S1 is collected and input into the target detection model in the following manner: fisheye cameras are arranged at four positions, front, back, left and right, of the vehicle body to obtain picture frames around the vehicle; the camera internal parameters are calibrated using the checkerboard calibration method in the ROS library and the initUndistortRectifyMap method in OpenCV is called to dedistort and splice the picture frames around the vehicle to obtain a corrected image; the corrected image is inversely transformed using the DLT algorithm based on the image homography matrix to obtain a four-way bird's-eye view; the four-way bird's-eye view is registered and spliced using the coordinate system transformation method to form a panoramic surround view; the panoramic surround view is fused using the weighted average method in the pixel-level fusion algorithm to obtain the final panoramic bird's-eye view as the surround monitoring image and input into the target detection model.
[0008] In another implementation, the target detection model is optimized by yolov3spp.
[0009] In another implementation, the pre-trained convolutional neural network of step S2 is used as the feature extraction network, the Darknet53 network is kept as the backbone network, the FPN part is replaced by BiFPN in the detection part, and multi-layer detection is performed by fusing feature information of different scales. The features of large detection targets are extracted from low-size feature dimensions, and the features of small detection targets are extracted from high-size feature dimensions with strong semantic meanings.
[0010] In another implementation, step S3 extracts the global information of the parking space and specifically includes: combining the entrance location, type and occupancy status in the global information to generate an intermediate parking space detection result; applying a 3×3×512 convolution filter, using a Sigmoid function to calculate the probability of each grid cell being located in the parking space, and outputting a 13×13×1 tensor, and the calculation formula is:
[0011]
[0012] in, is the true value of whether the center of the i-th grid unit is located in the parking space, is the network prediction value, Indicates whether the center of the i-th grid cell is located in the parking space, μ sl is the compensation coefficient;
[0013] Apply four 3×3×512 convolution filters and the sigmoid function to generate a 13×13×4 tensor to represent the relative position from the center of the grid cell to the two connection points of the parking space; apply three 3×3×512 convolution filters and the softmax function to generate a 13×13×3 tensor to represent whether the parking space is a vertical parking space, a parallel parking space, or a diagonal parking space; apply a 3×3×512 convolution filter and the sigmoid function to generate a 13×13×1 tensor to represent whether the parking space is occupied, and the calculation formula is:
[0014]
[0015] Among them, the actual value of the occupancy of the parking space at the center of the i-th grid unit is, is the network prediction value, Indicates whether the center of the i-th grid cell is located in an occupied parking space, Indicates whether the center of the i-th grid cell is located in an empty parking space, μ oc and μ va is the compensation coefficient.
[0016] In another implementation, the step S3 of extracting the local information of the parking space specifically includes: generating a detailed connection point detection result in combination with the position and direction of the connection point in the local information; using a 3×3×512 convolution filter and a sigmoid function to generate a 13×13×1 tensor to represent the probability of whether the grid unit contains a connection point, and the calculation formula is:
[0017]
[0018] in, is the true value of whether the i-th grid cell contains a connection point, is the network prediction value, Indicates whether the i-th grid cell contains a connection point, μ jn is the compensation coefficient;
[0019] Using two 3×3×512 convolution filters and a sigmoid function, a 13×13×2 tensor is generated to represent the relative position information from the center of the grid unit to the connection point. The calculation formula is:
[0020]
[0021] in, is the true relative position of the center of the i-th grid cell to the connection point contained in the cell, is the network prediction value, W ce and H ce is the width and height of the grid unit;
[0022] Using two 3×3×512 convolution filters and a sigmoid function, a 13×13×2 tensor is generated to represent the direction of the connection point. The calculation formula is:
[0023]
[0024] in, is the true direction vector of the connection point in the i-th grid unit, is the network prediction value.
[0025] In another implementation, the non-maximum suppression technology based on connection points adopted in step S4 specifically includes: non-maximum suppression technology based on connection points: the goal is to integrate global information and local information and select the most accurate connection points, including: a. extracting preliminary parking space connection points from global information and local information, wherein the connection points provided by the global information are based on the prediction of the overall parking space, and the connection points provided by the local information are based on the prediction of more refined local features; b. for each global connection point, find a local connection point within a preset distance; c. for each global connection point, find a local connection point within a preset distance; through the above steps ac, eliminate redundant global connection points and retain more accurate local connection points; based on parking spaces Non-maximum suppression technology: The goal is to further eliminate redundant parking space detection frames after integrating the connection point information and select the parking space with the highest confidence, including: d. Generate preliminary parking space detection frames based on the integrated connection point information, and sort the preliminary parking space detection frames according to the confidence of each preliminary parking space detection frame; e. Start with the preliminary parking space detection frame with the highest confidence, take it as the final result, and remove other preliminary parking space detection frames whose overlap with the preliminary parking space detection frame with the highest confidence exceeds a specified threshold; continue to select the preliminary parking space detection frame with the highest confidence from the remaining preliminary parking space detection frames, and repeat the above steps de until all preliminary parking space detection frames are processed.
[0026] According to a second aspect of an embodiment of the present invention, there is provided an information fusion parking space detection system based on deep learning, comprising: an image processing module, for dividing a collected 416×416 pixel surround monitoring image into a plurality of 13×13 grid units, and setting an 8-pixel overlapping area reserved between adjacent grid units; a feature extraction module, for performing multi-layer convolution and pooling processing on each grid unit through a pre-trained convolutional neural network to extract a feature map of size 13×13×512; an information extraction module, for applying a specified number of 3×3×512 convolution filters to the feature map, and extracting global information of the parking space through Sigmoid function calculation, wherein the parking space entrance position is obtained by calculating the pixel mean of the specified area in the feature map and the pre-trained convolutional neural network. A threshold comparison is determined, the type is determined by Softmax classification based on the pixel value distribution of different channels in the feature map, the occupancy status is judged according to the relationship between the pixel brightness of the specified area in the feature map and the set brightness threshold, and the local information of the parking space is extracted by a calculation method based on geometric relationships, including: determining the position of the connection point by calculating the coordinates of the specified key point in the feature map, and determining the direction by calculating the vector direction of adjacent key points; a result generation module is used for replacing the global connection point with the local connection point by a non-maximum suppression technology based on the connection point when the distance between the connection point in the global information and the connection point in the local information is less than 20 pixels, and integrating the global information and the local information of the parking space to generate the final parking space detection result.
[0027] According to a third aspect of an embodiment of the present invention, there is provided an electronic device, comprising a processor and a memory storing a program, wherein the program comprises instructions, and when the instructions are executed by the processor, the processor executes the steps executed by the method of the first aspect.
[0028] According to a fourth aspect of an embodiment of the present invention, there is provided a computer storage medium on which a computer program is stored. When the program is executed by a processor, the method of the first aspect described above is implemented.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] The method of the present invention first divides the image into grid units, extracts feature maps by convolutional neural networks, and obtains global information (including entrance location, type and occupancy) and local information (such as connection point location and direction) of parking spaces, and then integrates the information using non-maximum suppression technology based on connection points and parking spaces to obtain accurate detection results. The present invention improves the accuracy and efficiency of parking space detection, has significant advantages such as high detection rate, high precision and rapid detection, performs well under various environmental conditions, and has high practical value and broad application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0032] Figure 1 This is a flowchart of the steps of the information fusion parking space detection method based on deep learning of the present invention.
[0033] Figure 2 This is a schematic diagram of the network architecture of the information fusion parking space detection method based on deep learning of the present invention, showing the feature extraction network, information extractor and the connection relationship between them.
[0034] Figure 3 This is a schematic diagram of the global information extraction process of the information fusion parking space detection method based on deep learning of the present invention, which details the global information extraction steps and network structure.
[0035] Figure 4 It is a schematic diagram of the local information extraction process of the information fusion parking space detection method based on deep learning of the present invention, and details the local information extraction steps and network structure. DETAILED DESCRIPTION
[0036] In order to have a clearer understanding of the technical features, purposes and effects of the embodiments of the present invention, the specific implementation of the embodiments of the present invention is now described with reference to the accompanying drawings.
[0037] In this document, “exemplary” means “serving as an example, instance or illustration”, and any illustration or implementation described in this document as “exemplary” should not be construed as a more preferred or more advantageous technical solution.
[0038] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in the field based on the embodiments in the embodiments of the present invention should fall within the scope of protection of the embodiments of the present invention.
[0039] The specific implementation of the embodiment of the present invention is further described below in conjunction with the accompanying drawings of the embodiment of the present invention.
[0040] See also Figure 1-Figure 4 The information fusion parking space detection method based on deep learning provided by the present invention includes:
[0041] Step S1, dividing the collected 416×416 pixel surround view monitoring image (AVM) into a number of 13×13 grid units, and setting an 8-pixel overlapping area between adjacent grid units;
[0042] Step S2: Perform multi-layer convolution and pooling processing on each grid unit through a pre-trained convolutional neural network to extract a feature map of size 13×13×512;
[0043] Step S3, applying a specified number of 3×3×512 convolution filters to the feature map, and extracting the global information of the parking space through Sigmoid function calculation, wherein the parking space entrance position is determined by calculating the pixel mean of the specified area in the feature map and comparing it with a preset threshold, the type is determined by Softmax classification based on the pixel value distribution of different channels in the feature map, and the occupancy status is determined according to the relationship between the pixel brightness of the specified area in the feature map and the set brightness threshold, and at the same time, the local information of the parking space is extracted by a calculation method based on geometric relationships, including: determining the position of the connection point by calculating the coordinates of the specified key point in the feature map, and determining the direction by calculating the vector direction of adjacent key points;
[0044] Step S4: When the distance between the connection point in the global information and the connection point in the local information is less than 20 pixels, the connection point-based non-maximum suppression technology is used to replace the global connection point with the local connection point, and the global information and local information of the parking space are integrated to generate the final parking space detection result.
[0045] In one implementation, the surround monitoring image of step S1 is collected and input into the target detection model in the following manner: fisheye cameras are arranged at four positions, front, back, left and right, of the vehicle body to obtain picture frames around the vehicle; the camera internal parameters are calibrated using the checkerboard calibration method in the ROS library and the initUndistortRectifyMap method in OpenCV is called to dedistort and splice the picture frames around the vehicle to obtain a corrected image; the corrected image is inversely transformed using the DLT algorithm based on the image homography matrix to obtain a four-way bird's-eye view; the four-way bird's-eye view is registered and spliced using the coordinate system transformation method to form a panoramic surround view; the panoramic surround view is fused using the weighted average method in the pixel-level fusion algorithm to obtain the final panoramic bird's-eye view as the surround monitoring image and input into the target detection model.
[0046] In another implementation, the target detection model is optimized by yolov3spp.
[0047] In another implementation, the pre-trained convolutional neural network of step S2 is used as the feature extraction network, the Darknet53 network is kept as the backbone network, the FPN part is replaced by BiFPN in the detection part, and multi-layer detection is performed by fusing feature information of different scales. The features of large detection targets are extracted from low-size feature dimensions, and the features of small detection targets are extracted from high-size feature dimensions with strong semantic meanings.
[0048] In another implementation, step S3 extracts the global information of the parking space and specifically includes: combining the entrance location, type and occupancy status in the global information to generate an intermediate parking space detection result; applying a 3×3×512 convolution filter, using a Sigmoid function to calculate the probability of each grid cell being located in the parking space, and outputting a 13×13×1 tensor, and the calculation formula is:
[0049]
[0050] in, is the true value of whether the center of the i-th grid unit is located in the parking space, is the network prediction value, Indicates whether the center of the i-th grid cell is located in the parking space, μ sl is the compensation coefficient;
[0051] Apply four 3×3×512 convolution filters and the sigmoid function to generate a 13×13×4 tensor to represent the relative position from the center of the grid cell to the two connection points of the parking space; apply three 3×3×512 convolution filters and the softmax function to generate a 13×13×3 tensor to represent whether the parking space is a vertical parking space, a parallel parking space, or a diagonal parking space; apply a 3×3×512 convolution filter and the sigmoid function to generate a 13×13×1 tensor to represent whether the parking space is occupied, and the calculation formula is:
[0052]
[0053] Among them, the actual value of the occupancy of the parking space at the center of the i-th grid unit is, is the network prediction value, Indicates whether the center of the i-th grid cell is located in an occupied parking space, Indicates whether the center of the i-th grid cell is located in an empty parking space, μ oc and μ va is the compensation coefficient.
[0054] In another implementation, the step S3 of extracting the local information of the parking space specifically includes: generating a detailed connection point detection result in combination with the position and direction of the connection point in the local information; using a 3×3×512 convolution filter and a sigmoid function to generate a 13×13×1 tensor to represent the probability of whether the grid unit contains a connection point, and the calculation formula is:
[0055]
[0056] in, is the true value of whether the i-th grid cell contains a connection point, is the network prediction value, Indicates whether the i-th grid cell contains a connection point, μ jn is the compensation coefficient;
[0057] Using two 3×3×512 convolution filters and a sigmoid function, a 13×13×2 tensor is generated to represent the relative position information from the center of the grid unit to the connection point. The calculation formula is:
[0058]
[0059] in, is the true relative position of the center of the i-th grid cell to the connection point contained in the cell, is the network prediction value, W ce and H ce is the width and height of the grid unit;
[0060] Using two 3×3×512 convolution filters and a sigmoid function, a 13×13×2 tensor is generated to represent the direction of the connection point. The calculation formula is:
[0061]
[0062] in, is the true direction vector of the connection point in the i-th grid unit, is the network prediction value.
[0063] In another implementation, the connection point-based non-maximum suppression technique adopted in step S4 specifically includes:
[0064] Non-maximum suppression technology NMS based on connection points: The goal is to integrate global information and local information and select the most accurate connection points, including: a. Extracting preliminary parking space connection points from global information and local information, where the connection points provided by the global information are based on the prediction of the overall parking space, and the connection points provided by the local information are based on the prediction of more refined local features; b. For each global connection point, find a local connection point within a preset distance (such as 20 pixels); c. For each global connection point, find a local connection point within a preset distance (such as 20 pixels); Through the above steps ac, eliminate redundant global connection points and retain more accurate local connection points;
[0065] Non-maximum suppression technology NMS based on parking spaces: The goal is to further eliminate redundant parking space detection frames after integrating the connection point information, and select the parking space with the highest confidence, that is, the most likely parking space, including: d. Generate preliminary parking space detection frames based on the integrated connection point information, and sort the preliminary parking space detection frames according to the confidence (likelihood) of each preliminary parking space detection frame; e. Start with the preliminary parking space detection frame with the highest confidence, take it as the final result, and remove other preliminary parking space detection frames whose overlap (usually using IoU, intersection over union) with the preliminary parking space detection frame with the highest confidence exceeds the specified threshold; continue to select the preliminary parking space detection frame with the highest confidence from the remaining preliminary parking space detection frames, and repeat the above steps de until all preliminary parking space detection frames are processed.
[0066] Specifically, the scheme of the present invention is further described according to the following examples:
[0067] Step 1: Image acquisition and segmentation:
[0068] The AVM is acquired by fisheye cameras installed at four locations on the vehicle, front, back, left, and right. The fisheye cameras can capture a wide field of view around the vehicle and provide comprehensive environmental information for parking space detection.
[0069] The collected AVM image is divided into 13×13 grid cells, and an 8-pixel overlap area is reserved between adjacent grid cells. This division method helps to better capture the local features and global information of the parking space in the subsequent feature extraction and information extraction process.
[0070] The specific process of inputting the original image, i.e. the collected AVM image, is as follows:
[0071] First, fisheye cameras arranged around the vehicle body acquire image frames around the vehicle;
[0072] Next, the camera intrinsic parameters are calibrated using the checkerboard calibration method in the ROS library, and the initUndistortRectifyMap method in OpenCV is called to dedistort and stitch the original image frames.
[0073] Subsequently, the DLT algorithm based on the image homography matrix is used to perform inverse perspective transformation on the rectified image to obtain a bird’s-eye view of the four parts;
[0074] Then, the four bird's-eye views are registered and spliced through the coordinate system transformation method to form a panoramic ring view;
[0075] Finally, the weighted average method in the pixel-level fusion algorithm is used to fuse the panoramic ring view to obtain the final panoramic bird's-eye view and input it into the target detection model.
[0076] Step 2: Feature extraction:
[0077] Convolutional neural networks are used to extract features from each grid cell. The present invention attempts three backbone networks, namely Darknet53, ResNet50 and DenseNet121, which have been proven to have good performance in various applications.
[0078] When using these backbone networks, their fully connected layers are first removed, because fully connected layers usually introduce a large number of parameters, increase computational costs, and have relatively weak local feature extraction capabilities for images.
[0079] After removing the fully connected layers, the feature map dimensions obtained from these backbone networks are 13 × 13 × 512. This feature map contains rich feature information of each grid unit, providing a basis for subsequent information extraction.
[0080] The target detection model of the present invention is optimized by yolov3spp. The feature extraction network keeps the Darknet53 network as the backbone network. The detection part replaces the FPN part with BiFPN. Multi-layer detection is performed by fusing feature information of different scales. The features of large detection targets are extracted from low-size feature dimensions, and the features of small detection targets are extracted from high-size feature dimensions with strong semantic meanings. A bidirectional feature transfer path is added, and learnable weights are introduced to balance the importance of different input features.
[0081] Step 3: Information extraction:
[0082] (1) Combine global information (entrance location, type, occupancy) to generate intermediate parking space detection results.
[0083] Apply a 3×3×512 convolution filter and use the Sigmoid function to calculate the probability of each grid cell being in a parking space, outputting a 13×13×1 tensor calculated as:
[0084]
[0085] in, is the true value of whether the center of the i-th grid unit is located in the parking space (1 means it is located, 0 means it is not located), is the network prediction value, Indicates whether the center of the i-th grid unit is located in the parking space (1 means it is, 0 means it is not), μ sl is the compensation coefficient.
[0086] Apply four 3×3×512 convolutional filters and the sigmoid function to generate a 13×13×4 tensor representing the relative position from the center of the grid cell to the two connection points of the parking space. Apply three 3×3×512 convolutional filters and the softmax function to generate a 13×13×3 tensor representing whether the parking space is a vertical, parallel, or diagonal parking space.
[0087] Using a 3×3×512 convolution filter and a sigmoid function, a 13×13×1 tensor is generated to indicate whether the parking space is occupied. The calculation formula is:
[0088]
[0089] Among them, the actual occupancy value of the parking space at the center of the i-th grid unit (1 means occupied, 0 means free), is the network prediction value, Indicates whether the center of the i-th grid cell is located in an occupied parking space (1 for located, 0 for not located), Indicates whether the center of the i-th grid unit is located in an empty parking space (1 means it is, 0 means it is not), μ oc and μ va is the compensation coefficient.
[0090] (2) Combine local information (location and orientation of tie points) to generate detailed tie point detection results.
[0091] Using a 3×3×512 convolution filter and a sigmoid function, a 13×13×1 tensor is generated, which represents the probability of whether a grid cell contains a connection point. The calculation formula is:
[0092]
[0093] in, is the true value of whether the i-th grid cell contains the connection point (1 means it contains, 0 means it does not contain), is the network prediction value, Indicates whether the i-th grid cell contains a connection point (1 means it contains, 0 means it does not contain), μ jn is the compensation coefficient.
[0094] Using two 3×3×512 convolution filters and a sigmoid function, a 13×13×2 tensor is generated, which represents the relative position information from the center of the grid unit to the connection point. The calculation formula is:
[0095]
[0096] in, is the true relative position of the center of the i-th grid cell to the connection point contained in the cell, is the network prediction value, W ce and H ce is the width and height of the grid cell.
[0097] Using two 3×3×512 convolution filters and a sigmoid function, a 13×13×2 tensor is generated to represent the direction of the connection point. The calculation formula is:
[0098]
[0099] in, is the true direction vector of the connection point in the i-th grid unit, is the network prediction value.
[0100] In the global information, the parking space entrance possibility tensor represents the probability that the center of each grid cell is located in a parking space. The darker the color (such as green), the greater the probability, otherwise the lighter the color (such as gray). The parking space type tensor represents different types of parking spaces (vertical, parallel, inclined) by using different colors (such as blue, magenta, red). The parking space occupancy tensor uses different colors (such as purple for occupied and yellow for free) to represent the occupancy status of the parking space.
[0101] In the local information, the probability tensor of whether a grid cell contains a connection point is displayed in green, otherwise it is gray. The information of the relative position of the connection point is represented by a blue arrow, and the direction of the arrow points from the center of the grid cell to the connection point. The information of the direction of the connection point is represented by a red arrow, and the direction of the arrow represents the direction of the connection point.
[0102] Global information and local information complement each other. Global information provides an overall overview of parking spaces, while local information provides more accurate connection point location and direction information, which helps to improve the positioning accuracy of parking space detection.
[0103] Step 4: Overall framework and training:
[0104] (1) Feature extractor initialization:
[0105] The present invention uses weights pre-trained on ImageNet to initialize the feature extractor. ImageNet is a large-scale image dataset, and the models pre-trained on it can usually learn common image features. By using these pre-trained weights, the training process of the model can be accelerated and the performance of the model can be improved.
[0106] Specifically, for the three backbone networks VGG16, ResNet50, and DenseNet121, after removing their fully connected layers, the pre-trained weights are applied to the remaining network structure to initialize the parameters of the feature extractor. This enables the model to have a certain image understanding ability at the beginning of training, so that it can converge to better performance faster.
[0107] (2) Convolution filter initialization:
[0108] The 14 convolutional filters used to extract global and local information are initialized using the Xavier uniform initializer. The Xavier uniform initializer is a commonly used initialization method that can make the parameters of the convolutional filters have a reasonable distribution at the beginning, which helps to avoid the problem of gradient disappearance or explosion during training.
[0109] These convolution filters include 9 3×3×512 convolution filters for global information extraction and 5 3×3×512 convolution filters for local information extraction. Through reasonable initialization, these filters can better learn the relevant features of parking spaces at the beginning of training.
[0110] (3)Optimizer settings:
[0111] The Adam optimizer is used during the training process to optimize the parameters of the model. The Adam optimizer is an adaptive optimization algorithm that can automatically adjust the learning rate according to the changes in parameters during the training process, thereby improving the efficiency and stability of the training.
[0112] The parameters of the Adam optimizer are set to: learning rate 10^-4, β1 0.9, β2 0.999, ε 10^-8. The selection of these parameters has been verified by experiments and can achieve good results in different training scenarios.
[0113] (4) Training parameters:
[0114] The network is trained for 100 cycles with a batch size of 24. The selection of training cycles is to ensure that the model can fully learn the features in the data, while the batch size setting affects the efficiency of training and the generalization ability of the model. After many experiments, 100 cycles and a batch size of 24 have been proven to be more appropriate parameter settings.
[0115] The embodiment of the present invention further provides an information fusion parking space detection system based on deep learning, comprising:
[0116] An image processing module is used to divide the collected 416×416 pixel surround monitoring image into a number of 13×13 grid units, and set an 8-pixel overlapping area between adjacent grid units;
[0117] The feature extraction module is used to perform multi-layer convolution and pooling processing on each grid unit through a pre-trained convolutional neural network to extract a feature map of size 13×13×512;
[0118] An information extraction module is used to apply a specified number of 3×3×512 convolution filters to the feature map, and extract the global information of the parking space through Sigmoid function calculation, wherein the parking space entrance position is determined by calculating the pixel mean of the specified area in the feature map and comparing it with a preset threshold, the type is determined by Softmax classification based on the pixel value distribution of different channels in the feature map, and the occupancy status is determined based on the relationship between the pixel brightness of the specified area in the feature map and the set brightness threshold, and the local information of the parking space is extracted by using a calculation method based on geometric relationships, including: determining the position of the connection point by calculating the coordinates of the specified key point in the feature map, and determining the direction by calculating the vector direction of adjacent key points;
[0119] The result generation module is used to replace the global connection point with the local connection point when the distance between the connection point in the global information and the connection point in the local information is less than 20 pixels, and integrate the global information and local information of the parking space to generate the final parking space detection result.
[0120] In summary, the present invention first divides the image into grid units, extracts feature maps by convolutional neural networks, and obtains global information (including entrance location, type and occupancy) and local information (such as connection point location and direction) of parking spaces, and then integrates the information using non-maximum suppression technology based on connection points and parking spaces to obtain accurate detection results. The present invention improves the accuracy and efficiency of parking space detection, and has significant advantages such as high detection rate, high precision and rapid detection. It performs well under various environmental conditions and has high practical value and broad application prospects.
[0121] As another example, the present invention also provides an electronic device, and now will describe an electronic device that can be used as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic devices are intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples, and are not intended to limit the implementation of the present invention described and / or required herein.
[0122] The electronic device may include: a processor (processor), a communication interface (CommunicationsInterface), a memory (memory) and a communication bus.
[0123] The processor, the communication interface and the memory communicate with each other through the communication bus. The communication interface is used to communicate with other electronic devices or servers.
[0124] The processor is used to execute the program, and specifically can execute the relevant steps in the above method embodiment.
[0125] Specifically, the program may include program codes including computer operation instructions.
[0126] The processor may be a CPU, or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. The one or more processors included in the smart device may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.
[0127] The memory is used to store programs and may include a high-speed RAM memory and may also include a non-volatile memory, such as at least one disk memory.
[0128] When the program is executed by the processor, it is used to enable the electronic device to execute the information fusion parking space detection method based on deep learning of the present invention.
[0129] In addition, the specific implementation of each step in the program can refer to the corresponding description of the corresponding steps and units in the above method embodiment, which will not be repeated here. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described devices and modules can refer to the corresponding process description in the above method embodiment, which will not be repeated here.
[0130] The exemplary embodiments of the present invention further provide a computer storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the methods of the various embodiments of the present invention. The corresponding process descriptions in the aforementioned method embodiments may be referred to and will not be repeated here.
[0131] The above-described method according to an embodiment of the present invention may be implemented in hardware, firmware, or as software or computer code that may be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code that is originally stored in a remote recording medium or a non-temporary machine-readable medium downloaded over a network and will be stored in a local recording medium, so that the method described herein may be stored in such software processing on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that a computer, processor, microprocessor controller, or programmable hardware includes a storage component (e.g., RAM, ROM, flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by a computer, processor, or hardware, the method described herein is implemented. In addition, when a general-purpose computer accesses the code for implementing the method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the method shown herein.
[0132] Thus far, specific embodiments of the present invention have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired results. Additionally, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing may be advantageous.
[0133] It should be understood that although this specification is described according to various embodiments, not every embodiment contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.
[0134] Finally, it should be noted that the above implementation methods are only used to illustrate the embodiments of the present invention, and are not limitations of the embodiments of the present invention. Ordinary technicians in the relevant technical field can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present invention. The patent protection scope of the embodiments of the present invention should be defined by the claims.
Claims
1. A parking space detection method based on deep learning information fusion, characterized in that: include: Step S1, dividing the collected 416×416 pixel surround monitoring image into a number of 13×13 grid units, and setting an 8-pixel overlapping area between adjacent grid units; Step S2: Perform multi-layer convolution and pooling processing on each grid unit through a pre-trained convolutional neural network to extract a feature map of size 13×13×512; Step S3, applying a specified number of 3×3×512 convolution filters to the feature map, and extracting the global information of the parking space through Sigmoid function calculation, wherein the parking space entrance position is determined by calculating the pixel mean of the specified area in the feature map and comparing it with a preset threshold, the type is determined by Softmax classification based on the pixel value distribution of different channels in the feature map, and the occupancy status is determined according to the relationship between the pixel brightness of the specified area in the feature map and the set brightness threshold, and at the same time, the local information of the parking space is extracted by a calculation method based on geometric relationships, including: determining the position of the connection point by calculating the coordinates of the specified key point in the feature map, and determining the direction by calculating the vector direction of adjacent key points; Step S4: When the distance between the connection point in the global information and the connection point in the local information is less than 20 pixels, the connection point-based non-maximum suppression technology is used to replace the global connection point with the local connection point, and the global information and local information of the parking space are integrated to generate the final parking space detection result.
2. The method according to claim 1, characterized in that: The surround monitoring image of step S1 is collected and input into the target detection model in the following manner: The image frames around the vehicle are obtained by using fisheye cameras arranged at four positions, front, back, left, and right of the vehicle body; Use the checkerboard calibration method in the ROS library to calibrate the camera internal parameters and call the initUndistortRectifyMap method in OpenCV to dedistort and stitch the image frames around the vehicle to obtain the corrected image; The DLT algorithm based on the image homography matrix is used to perform inverse perspective transformation on the rectified image to obtain a four-way bird's-eye view; The four bird's-eye views are registered and spliced through the coordinate system transformation method to form a panoramic ring view; The weighted average method in the pixel-level fusion algorithm is used to fuse the panoramic surround view, and the final panoramic bird's-eye view image is obtained as the surround view monitoring image and input into the target detection model.
3. The method according to claim 2, characterized in that The target detection model is optimized by yolov3spp.
4. The method according to claim 1, characterized in that The pre-trained convolutional neural network of step S2 is used as the feature extraction network, the Darknet53 network is kept as the backbone network, the FPN part is replaced by BiFPN in the detection part, and multi-layer detection is performed by means of fusing feature information of different scales, extracting features of large detection targets from low-scale feature dimensions, and extracting features of small detection targets from high-scale feature dimensions with strong semantic meanings.
5. The method according to claim 1, characterized in that The step S3 of extracting the global information of the parking space specifically includes: Combine the entrance location, type and occupancy in the global information to generate the intermediate parking space detection results; Apply a 3×3×512 convolution filter and use the Sigmoid function to calculate the probability of each grid cell being in a parking space, outputting a 13×13×1 tensor calculated as: in, is the true value of whether the center of the i-th grid unit is located in the parking space, is the network prediction value, Indicates whether the center of the i-th grid cell is located in the parking space, μ sl is the compensation coefficient; Apply four 3×3×512 convolutional filters and the sigmoid function to generate a 13×13×4 tensor representing the relative position from the center of the grid cell to the two connection points of the parking space; Apply three 3×3×512 convolutional filters and a softmax function to generate a 13×13×3 tensor to indicate whether the parking space is a perpendicular parking space, a parallel parking space, or a diagonal parking space; Apply a 3×3×512 convolution filter and a sigmoid function to generate a 13×13×1 tensor to indicate whether the parking space is occupied. The calculation formula is: Among them, the actual value of the occupancy of the parking space at the center of the i-th grid unit is, is the network prediction value, Indicates whether the center of the i-th grid cell is located in an occupied parking space, Indicates whether the center of the i-th grid cell is located in an empty parking space, μ oc and μ va is the compensation coefficient.
6. The method according to claim 1, characterized in that The step S3 of extracting the local information of the parking space specifically includes: Generate detailed connection point detection results by combining the location and direction of the connection points in the local information; Using a 3×3×512 convolution filter and a sigmoid function, a 13×13×1 tensor is generated to represent the probability of whether a grid cell contains a connection point. The calculation formula is: in, is the true value of whether the i-th grid cell contains a connection point, is the network prediction value, Indicates whether the i-th grid cell contains a connection point, μ jn is the compensation coefficient; Using two 3×3×512 convolution filters and a sigmoid function, a 13×13×2 tensor is generated to represent the relative position information from the center of the grid unit to the connection point. The calculation formula is: in, is the true relative position of the center of the i-th grid cell to the connection point contained in the cell, is the network prediction value, W ce and H ce is the width and height of the grid unit; Using two 3×3×512 convolution filters and a sigmoid function, a 13×13×2 tensor is generated to represent the direction of the connection point. The calculation formula is: in, is the true direction vector of the connection point in the i-th grid unit, is the network prediction value.
7. The method according to claim 1, characterized in that The non-maximum suppression technology based on connection points adopted in step S4 specifically includes: Non-maximum suppression technology based on connection points: The goal is to integrate global information and local information and select the most accurate connection points, including: a. Extracting preliminary parking space connection points from global information and local information, where the connection points provided by global information are based on the prediction of the overall parking space, while the connection points provided by local information are based on the prediction of more refined local features; b. For each global connection point, find the local connection point within the preset distance; c. For each global connection point, find the local connection point within the preset distance; Through the above steps ac, redundant global connection points are eliminated and more accurate local connection points are retained; Parking space-based non-maximum suppression technology: The goal is to further eliminate redundant parking space detection boxes after integrating the connection point information and select the parking space with the highest confidence, including: d. Generate preliminary parking space detection frames according to the integrated connection point information, and sort the preliminary parking space detection frames according to the confidence of each preliminary parking space detection frame; e. Starting from the preliminary parking space detection frame with the highest confidence, taking it as the final result, and removing other preliminary parking space detection frames whose overlap with the preliminary parking space detection frame with the highest confidence exceeds a specified threshold; Continue to select the preliminary parking space detection frame with the highest confidence from the remaining preliminary parking space detection frames, and repeat the above steps until all preliminary parking space detection frames are processed.
8. An information fusion parking space detection system based on deep learning, characterized in that: include: An image processing module is used to divide the collected 416×416 pixel surround monitoring image into a number of 13×13 grid units, and set an 8-pixel overlapping area between adjacent grid units; The feature extraction module is used to perform multi-layer convolution and pooling processing on each grid unit through a pre-trained convolutional neural network to extract a feature map of size 13×13×512; An information extraction module is used to apply a specified number of 3×3×512 convolution filters to the feature map, and extract the global information of the parking space through Sigmoid function calculation, wherein the parking space entrance position is determined by calculating the pixel mean of the specified area in the feature map and comparing it with a preset threshold, the type is determined by Softmax classification based on the pixel value distribution of different channels in the feature map, and the occupancy status is determined based on the relationship between the pixel brightness of the specified area in the feature map and the set brightness threshold, and the local information of the parking space is extracted by using a calculation method based on geometric relationships, including: determining the position of the connection point by calculating the coordinates of the specified key point in the feature map, and determining the direction by calculating the vector direction of adjacent key points; The result generation module is used to replace the global connection point with the local connection point when the distance between the connection point in the global information and the connection point in the local information is less than 20 pixels, and integrate the global information and local information of the parking space to generate the final parking space detection result.
9. An electronic device, characterized in that: include: processor; A memory for storing programs; The program includes instructions, which, when executed by the processor, cause the processor to perform the steps of the method as claimed in any one of claims 1 to 7.
10. A computer storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Unmanned aerial vehicle detection method based on improved YOLOX algorithm
CN116452994A
Method and device for detecting parking space, storage medium and vehicle
CN119068448A