Vision-based parking lot vehicle management method and device
By using a combination of deep learning image classification model and parking space sensors in the parking lot, the parking lot vehicle status is monitored and managed in real time, and the traditional management methods are solved, achieving more efficient and accurate parking lot management.
Patent Information
- Application Number
- CN202510118856.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional manual patrols and manual recordings cannot meet the rapidly growing urban traffic demand and real-time requirements of parking management, resulting in less accurate, inefficient and costly parking lot vehicle management.
The vision-based parking lot vehicle management method is adopted. By inputting the parking space image of each parking space in the parking lot into the deep learning image classification model for processing, the parking space status probability is generated, and the parking space status information obtained by the parking space sensor is weighted to determine the real-time status of each parking space and feedback it to the vehicles entering the parking lot in real time so that they can quickly find the spare parking space.
It improves the accuracy and efficiency of parking lot vehicle management, reduces management costs, and realizes real-time monitoring and management of parking lot vehicles.
Smart Images

Figure CN119942838A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of parking lot management based on visual recognition, and in particular to a parking lot vehicle management method and device based on vision. Background Art
[0002] In modern urban life, driving has become a means of transportation for many people. However, due to the large number of vehicles, parking spaces in parking lots are in short supply, and efficient use of parking spaces has become an important challenge. Especially in some small spaces such as parks or underground parking lots with high demand for parking spaces, the distribution of parking spaces is complex, the daily vehicle flow fluctuates greatly, and there are cases of random parking and random occupation of parking spaces. It is particularly important to manage the efficiency of parking space use and understand the parking situation in real time.
[0003] Traditional manual inspections and manual records have problems such as low accuracy, slow efficiency, and high cost, and cannot meet the rapidly growing urban traffic needs and real-time requirements of parking management. Summary of the invention
[0004] The purpose of the present invention is to provide a vision-based parking lot vehicle management method and device to alleviate the technical problems of low accuracy, low efficiency and high cost of parking lot vehicle management.
[0005] In a first aspect, an embodiment of the present invention provides a method for managing parking lots based on vision, comprising:
[0006] Inputting a parking space image of each parking space in the parking lot into a deep learning image classification model for processing, and outputting a parking space state probability of each parking space; wherein the deep learning image classification model includes a residual block structure for generating a multi-scale feature map;
[0007] Performing weighted processing based on the parking space state probability and the parking space state information of each parking space obtained from the parking space sensor to determine the real-time state of each parking space;
[0008] The vehicles entering and exiting the parking lot are monitored, and the real-time status of each parking space is fed back to the vehicles entering the parking lot, so that the vehicles can park in a vacant parking space based on the real-time status of each parking space.
[0009] Furthermore, the step of inputting the parking space image of each parking space in the parking lot into the deep learning image classification model for processing and outputting the parking space state probability of each parking space includes:
[0010] Obtain a parking space image of each parking space in the parking lot, and input the parking space image into a deep learning image classification model; wherein the deep learning image classification model includes a first-stage module, a second-stage module, a third-stage module, a fourth-stage module, a fifth-stage module, a global average pooling layer and a fully connected layer connected in sequence; the first-stage module includes a 7×7 convolutional layer and a 3×3 maximum pooling layer; the second-stage module consists of three residual block structures, the third-stage module consists of four residual block structures, and the fourth-stage module consists of six residual block structures;
[0011] Based on the processing of the residual block structure in the deep learning image classification model, a multi-scale feature map of each parking space is obtained;
[0012] Each of the multi-scale feature maps is processed by the remaining structures in the deep learning image classification model to output the parking space state probability of each of the parking spaces.
[0013] Furthermore, based on the processing of the residual block structure in the deep learning image classification model, the step of obtaining a multi-scale feature map of each parking space includes:
[0014] The feature map of each residual block structure is first processed by a 1×1 convolution layer, and then processed in parallel by a 1×1 convolution layer, a 3×3 convolution layer, and a 5×5 convolution layer to obtain feature maps of three scales;
[0015] The feature maps of the three scales are concatenated in the channel dimension, and then input into a 1×1 convolutional layer for compression to obtain a multi-scale feature map of each parking space.
[0016] Furthermore, the step of performing weighted processing based on the parking space state probability and the parking space state information of each parking space obtained from the parking space sensor to determine the real-time state of each parking space includes:
[0017] Acquiring parking space status information of each parking space from a parking space sensor provided at each parking space in the parking lot; wherein the parking space status information is used to indicate whether the parking space detected by the parking space sensor is a vacant parking space or a non-vacant parking space;
[0018] The parking space status probability and the parking space status information of each parking space are weightedly calculated according to preset weights to determine the real-time status of each parking space; wherein the parking space status probability includes the probability that the parking space is a vacant parking space and the probability that the parking space is a non-vacant parking space.
[0019] Furthermore, the step of monitoring vehicles entering and exiting the parking lot and feeding back the real-time status of each parking space to the vehicles entering the parking lot includes:
[0020] Identify the license plate information of each vehicle entering and leaving the parking lot, and record the entry time and exit time of each vehicle;
[0021] Feedback the real-time status of each parking space to each vehicle entering the parking lot;
[0022] Inputting the surveillance video of the parking lot into the deep learning target detection model for processing to determine the parking position of each vehicle;
[0023] The license plate information of each vehicle is bound to the parking location.
[0024] Furthermore, the step of identifying the license plate information of each vehicle entering and leaving the parking lot includes:
[0025] Acquire the license plate image of each vehicle entering and exiting the parking lot from the entrance and exit cameras of the parking lot;
[0026] Inputting the license plate image of each of the vehicles into a UNet image segmentation model, and outputting a first feature map, a second feature map, a third feature map, a fourth feature map, and a fifth feature map of each of the vehicles through downsampling branch processing of the UNet image segmentation model;
[0027] The first feature map is processed by a first feature extraction RFB module to obtain a first effective feature map; the second feature map is processed by a second feature extraction RFB module to obtain a second effective feature map; the third feature map is processed by a third feature extraction RFB module to obtain a third effective feature map; the fourth feature map is processed by a fourth feature extraction RFB module to obtain a fourth effective feature map; the fifth feature map is processed by a fifth feature extraction RFB module to obtain a fifth effective feature map;
[0028] Performing maximum pooling layer processing on the first effective feature map, and then performing channel splicing on the second effective feature map to obtain a first spliced feature map;
[0029] Performing maximum pooling layer processing on the first spliced feature map, and then performing channel splicing with the third effective feature map to obtain a second spliced feature map;
[0030] Performing maximum pooling layer processing on the second spliced feature map, and then performing channel splicing with the fourth effective feature map to obtain a third spliced feature map;
[0031] Performing maximum pooling layer processing on the third spliced feature map, and then performing channel splicing with the fifth effective feature map to obtain a fourth spliced feature map;
[0032] Compressing the fourth concatenated feature map through a 1×1 convolutional layer to output a final feature map of each vehicle;
[0033] The final feature map of each vehicle is input into the channel attention module for processing, and then optical character recognition is performed to obtain the license plate information of each vehicle.
[0034] Furthermore, the step of inputting the monitoring video of the parking lot into the deep learning target detection model for processing to determine the parking position of each vehicle includes:
[0035] Based on the surveillance video of the parking lot, obtaining a video frame at each moment;
[0036] Input the video frame into a deep learning target detection YOLOv11 model to extract the target features of each vehicle at each moment;
[0037] The target features at all adjacent moments are compared for similarity to obtain the driving trajectory of each vehicle, and the parking position of each vehicle is updated.
[0038] In a second aspect, an embodiment of the present invention further provides a parking lot vehicle management device based on vision, comprising:
[0039] A first processing module inputs a parking space image of each parking space in the parking lot into a deep learning image classification model for processing, and outputs a parking space state probability of each parking space; wherein the deep learning image classification model includes a residual block structure for generating a multi-scale feature map;
[0040] A second processing module performs weighted processing based on the parking space state probability and the parking space state information of each parking space obtained from the parking space sensor to determine the real-time state of each parking space;
[0041] The parking management module monitors the vehicles entering and exiting the parking lot, and feeds back the real-time status of each parking space to the vehicles entering the parking lot, so that the vehicles can park in a vacant parking space based on the real-time status of each parking space.
[0042] In a third aspect, an embodiment provides an electronic device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps of the method described in any of the aforementioned implementation methods are implemented.
[0043] In a fourth aspect, an embodiment provides a machine-readable storage medium, wherein the machine-readable storage medium stores machine-executable instructions. When the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the steps of the method described in any one of the aforementioned implementations.
[0044] The embodiment of the present invention provides a vision-based parking lot vehicle management method and device. First, a parking space map of each parking space in the parking lot is input into a deep learning image classification model. The residual block structure in the deep learning image classification model can be used to generate a multi-scale feature map. On this basis, a larger receptive field is provided to facilitate the deep learning image classification model to more accurately obtain the parking space status probability of each parking space. In order to further ensure the accuracy of parking space status recognition and avoid recognition errors caused by vehicle occlusion, light and other problems, the parking space status probability of each parking space is combined with the parking space status information collected by the parking space sensor for weighted calculation to obtain the real-time status of each parking space. The real-time status of each parking space is sent to the vehicle entering the parking lot, so that the vehicle entering the parking lot can quickly, accurately and conveniently drive to the vacant parking space for parking. At the same time, the vehicles entering and leaving the parking lot are monitored and recorded to achieve management of each vehicle in the parking lot.
[0045] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description and the drawings.
[0046] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0048] Figure 1 A flow chart of a vision-based parking lot vehicle management method provided by an embodiment of the present invention;
[0049] Figure 2 A schematic diagram of a residual block structure provided by an embodiment of the present invention;
[0050] Figure 3 A schematic diagram of a cross-layer fusion module structure provided by an embodiment of the present invention;
[0051] Figure 4 A schematic diagram of the structure of a channel attention module provided by an embodiment of the present invention;
[0052] Figure 5 A schematic diagram of functional modules of a vision-based parking lot vehicle management method and device provided by an embodiment of the present invention;
[0053] Figure 6 A schematic diagram of the hardware architecture of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0054] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0055] The embodiments of the present invention provide a vision-based parking lot vehicle management method and device, which can manage the parking of vehicles in the parking lot with low cost, high precision and high efficiency.
[0056] To facilitate understanding of this embodiment, a vision-based parking lot vehicle management method disclosed in an embodiment of the present invention is first introduced in detail. The method is applied to intelligent control devices such as servers, controllers, and host computers.
[0057] Figure 1 A flow chart of a vision-based parking lot vehicle management method provided by an embodiment of the present invention.
[0058] Reference Figure 1 , the method comprises the following steps:
[0059] Step S102: input the parking space image of each parking space in the parking lot into the deep learning image classification model for processing, and output the parking space state probability of each parking space.
[0060] Among them, the deep learning image classification model includes a residual block structure for generating a multi-scale feature map; based on the multi-scale feature map generated by the residual block structure, the deep learning image classification model can achieve a more accurate parking space status probability output.
[0061] Step S104 , performing weighted processing based on the parking space state probability and the parking space state information of each parking space obtained from the parking space sensor to determine the real-time state of each parking space.
[0062] Here, in order to further enhance the accuracy of the obtained parking space status, based on the above-mentioned embodiment, weighted processing is also performed in combination with the parking space status information collected by the parking space sensor.
[0063] Step S106, monitoring vehicles entering and exiting the parking lot, and feeding back the real-time status of each parking space to the vehicles entering the parking lot, so that the vehicles park in a vacant parking space based on the real-time status of each parking space.
[0064] It is understandable that parking instructions are provided to vehicles entering the parking lot based on a more accurate parking space status, while vehicles entering and leaving the parking lot are monitored to facilitate parking management.
[0065] In a preferred embodiment of practical application, the parking space map of each parking space in the parking lot is first input into the deep learning image classification model. The residual block structure in the deep learning image classification model can be used to generate a multi-scale feature map, on this basis, a larger receptive field is provided to facilitate the deep learning image classification model to more accurately obtain the parking space status probability of each parking space; in order to further ensure the accuracy of parking space status recognition and avoid recognition errors caused by vehicle occlusion, light and other problems, the parking space status probability of each parking space is combined with the parking space status information collected by the parking space sensor for weighted calculation to obtain the real-time status of each parking space; the real-time status of each parking space is sent to the vehicles entering the parking lot, so that the vehicles entering the parking lot can quickly, accurately and conveniently drive to the vacant parking spaces for parking. At the same time, the vehicles entering and leaving the parking lot are monitored and recorded to achieve management of each vehicle in the parking lot.
[0066] In some embodiments, step S102, the step of determining the parking space status by detecting and identifying the parking space image, includes:
[0067] Step 1.1), obtain the parking space image of each parking space in the parking lot, and input the parking space image into the deep learning image classification model.
[0068] Wherein, each parking space in the parking lot is provided with a parking space image acquisition device, such as a camera, which can provide a real-time parking space video stream or an offline parking space image.
[0069] Among them, the deep learning image classification model adopts the BoTNet-50 algorithm, including 5 stages (Stage 0, 1, 2, 3, 4) connected in sequence, such as the first stage module, the second stage module, the third stage module, the fourth stage module, the fifth stage module, and the global average pooling layer and the fully connected layer; the first stage module Stage 0 includes a 7×7 convolution layer and a 3×3 maximum pooling layer; the second stage module Stage 1 consists of three residual block structures, the third stage module Stage 2 consists of four residual block structures, and the fourth stage module Stage 3 consists of six residual block structures; the fifth stage module Stage 4 consists of 3 MHSA (Multi Head Self Attention) residual structures, and the MHSA residual structure consists of 2 1×1 convolutions, MHSA structure and jump connection.
[0070] Step 1.2), based on the processing of the residual block structure in the deep learning image classification model, a multi-scale feature map of each parking space is obtained.
[0071] For example, Figure 2 As shown in the figure, the residual block structure in the deep learning image classification model is composed of a 1×1 convolutional layer, a parallel 1×1 convolutional layer, a 3×3 convolutional layer and a 5×5 convolutional layer, a 1×1 convolutional layer and a skip connection. The residual block structure can generate a multi-scale feature map through the following steps, including:
[0072] In step 1.2.1), the feature map of each residual block structure is first processed by a 1×1 convolution layer, and then processed in parallel by a 1×1 convolution layer, a 3×3 convolution layer, and a 5×5 convolution layer to obtain feature maps of three scales.
[0073] It should be noted that the above-mentioned deep learning image classification model includes various modules of the residual block structure, and the feature maps of the input residual block structure are all processed according to steps 1.2.1) to 1.2.2).
[0074] In step 1.2.2), the feature maps of the three scales are concatenated in the channel dimension and then input into a 1×1 convolutional layer for compression to obtain a multi-scale feature map of each parking space.
[0075] The feature maps extracted by the three convolutional layers are spliced in the channel dimension, and then the spliced feature maps are input into a 1×1 convolution compression channel to obtain a multi-scale feature map. Compared with the feature maps extracted by traditional methods in the prior art, the multi-scale feature map has a larger receptive field.
[0076] Assuming that the feature map of the input residual block structure is x∈R(W×H×C), where W, H, and C represent the width, height, and number of channels of the feature map, respectively, the output multi-scale feature map y of the improved residual block structure is:
[0077] y=K1(C(K1(K1(x)),K2(K1(x)),K3(K1(x))))+x
[0078] In the formula, K1(·), K2(·), and K3(·) represent 1×1 convolution, 3×3 convolution, and 5×5 convolution, respectively, and C(·) represents channel concatenation.
[0079] Step 1.3), each multi-scale feature map is processed by the remaining structures in the deep learning image classification model, and the parking state probability of each parking space is output.
[0080] Here, the deep learning image classification model uses multi-scale feature maps to further accurately identify the parking space status probability, avoid the impact of parking space occlusion, and thus alleviate the problem of false detection caused by some parking spaces being occluded.
[0081] On the basis of the above-mentioned embodiment, in order to further improve the accuracy of parking space status recognition, step S104 performs weighted processing on the results of the parking space recognition module and the parking space sensor acquisition device, thereby finally determining whether there is an empty parking space, which specifically includes:
[0082] Step 2.1), obtaining parking space status information of each parking space from a parking space sensor installed at each parking space in the parking lot.
[0083] Among them, the parking space status information is used to characterize whether the parking space detected by the parking space sensor is a vacant parking space or a non-vacant parking space; the parking space sensor can be a geomagnetic parking space detector with a waterproof and pressure-resistant shell, which can be used in indoor and outdoor environments to detect parking space status information.
[0084] Step 2.2), the parking space status probability and parking space status information of each parking space are weightedly calculated according to preset weights to determine the real-time status of each parking space.
[0085] The parking space status probability includes the probability that the parking space is a vacant parking space and the probability that the parking space is a non-vacant parking space.
[0086] In order to further alleviate the problems of missed detection and false detection of parking space status due to irregular parking and partial obstruction of parking spaces, the embodiments of the present invention perform weighted processing on the parking space status probability calculated in the above embodiments and the parking space status information collected by the parking space sensor to determine the final parking space status.
[0087] Assume that the weights of the parking space status information collected by the parking space sensor and the parking space status probability calculated by the image input deep learning image classification model are W and 传感器 and W 算法 , if the parking space status information of the parking space sensor is a non-vacant parking space, the parking space status probability calculated by the image input deep learning image classification model is a vacant parking space and a non-vacant parking space, respectively. 空余 and P 非空余 , then the weighted probability of an empty parking space is (W 算法 ×P 空余 ), the weighted probability of a non-vacant parking space is (W 算法 ×P 非空余 +1×W 传感器 ), and the final parking space status is determined by comparing the weighted probability of vacant parking spaces with the probability of non-vacant parking spaces.
[0088] For example, the weights of the parking space status information collected by the parking sensor and the parking space status probability calculated by the image input deep learning image classification model are 0.4 and 0.6 respectively. If the parking space status information collected by the parking sensor is an available parking space, the parking space status probability calculated by the image input deep learning image classification model: the probability of a non-available parking space is 0.7, and the probability of an available parking space is 0.3. Then, after weighted processing in the above steps, the probability of a non-available parking space is 0.6*0.7=0.42, and the probability of an available parking space is 0.6*0.3+0.4*1=0.58. The probability of an available parking space is greater than the probability of a non-available parking space. Therefore, the final parking space status is determined to be an available parking space.
[0089] In some embodiments, the parking lot usage can be obtained by statistically analyzing the parking space status, vehicle license plate information and parking location, so as to realize intelligent management of the parking lot and help managers optimize parking resource allocation and improve service efficiency. Step S106 can be specifically realized by the following steps, including:
[0090] Step 3.1), identify the license plate information of each vehicle entering and leaving the parking lot, and record the entry time and exit time of each vehicle.
[0091] Exemplarily, the step of identifying the license plate information of the vehicle entering or leaving the vehicle according to the license plate image may include:
[0092] Step 3.1.1), obtain the license plate image of each vehicle entering and leaving the parking lot from the entrance and exit cameras of the parking lot.
[0093] The entrance and exit cameras of the parking lot can provide real-time video streams to obtain the license plate images of each vehicle entering and leaving.
[0094] In step 3.1.2), the license plate image of each vehicle is input into the UNet image segmentation model, and the first feature map, second feature map, third feature map, fourth feature map and fifth feature map of each vehicle are output through the downsampling branch of the UNet image segmentation model, and then input into the cross-layer fusion module.
[0095] like Figure 3 As shown, the first feature map, the second feature map, the third feature map, the fourth feature map and the fifth feature map are respectively input into the Receptive Field Block (RFB) module for feature extraction of the cross-layer fusion module in sequence from top to bottom.
[0096] Step 3.1.3), the first feature map is processed by the first feature extraction RFB module to obtain a first effective feature map; the second feature map is processed by the second feature extraction RFB module to obtain a second effective feature map; the third feature map is processed by the third feature extraction RFB module to obtain a third effective feature map; the fourth feature map is processed by the fourth feature extraction RFB module to obtain a fourth effective feature map; the fifth feature map is processed by the fifth feature extraction RFB module to obtain a fifth effective feature map.
[0097] Step 3.1.4), the first effective feature map is processed by the maximum pooling layer MP, and then channel-joined with the second effective feature map to obtain a first joint feature map.
[0098] Step 3.1.5), the first spliced feature map is processed by the maximum pooling layer MP, and then channel-spliced with the third effective feature map to obtain a second spliced feature map.
[0099] Step 3.1.6), the second spliced feature map is processed by the maximum pooling layer MP, and then channel-spliced with the fourth effective feature map to obtain the third spliced feature map.
[0100] Step 3.1.7), the third spliced feature map is processed by the maximum pooling layer MP, and then channel-joined with the fifth effective feature map to obtain the fourth spliced feature map.
[0101] In step 3.1.8), the fourth concatenated feature map is compressed through a 1×1 convolutional layer to output the final feature map of each vehicle.
[0102] In step 3.1.9), the final feature map of each vehicle is input into the channel attention module for processing, and then optical character recognition (OCR) is performed to obtain the license plate information of each vehicle.
[0103] It should be noted that after the license plate image is obtained by the parking lot entrance and exit camera, the license plate area is further accurately located through the deep learning image segmentation model, and then the license plate information in the license plate area is recognized using OCR technology. The image segmentation model can use the UNet algorithm. The structure of UNet is similar to the letter "U" and consists of a symmetrical encoder-decoder architecture. The encoder is responsible for capturing the features of the image, while the decoder converts these features into a segmented image of the same size as the input image.
[0104] In order to enable the UNet algorithm to more accurately segment the license plate area, the five feature maps output by the downsampling branch of the UNet image segmentation model are input into the cross-layer fusion module. The cross-layer fusion module is mainly composed of the RFB module. The five feature maps extracted by the downsampling branch are processed by the RFB module to obtain five valid feature maps; then the first valid feature map is subjected to maximum pooling and channel splicing with the second valid feature map, and the first spliced feature map is subjected to maximum pooling and channel splicing with the third valid feature map. Similarly, the final spliced feature map is passed through a 1×1 convolution compression channel to obtain the final feature map, thereby enhancing the network's feature extraction capabilities, which helps to retain image details and more accurately segment. In addition, the embodiment of the present invention also adds a channel attention mechanism on this basis, such as Figure 4 As shown in the figure, after the above-mentioned final feature map is input into the channel attention module, the feature size is first compressed from H×W×C to 1×1×C through the global average pooling GAP layer, and then the input features are recalibrated in the channel dimension through two fully connected layers FC and Sigmoid activation function to complete the feature adjustment. The channel attention module can improve the ability to extract small target features, highlight important feature information in the image, and thus improve the segmentation accuracy of the algorithm. On this basis, the segmented image is processed by OCR recognition to obtain more accurate license plate information.
[0105] Step 3.2), the real-time status of each parking space is fed back to each vehicle entering the parking lot.
[0106] Here, the real-time information of each parking space in the current parking lot can be sent to the vehicle computer of the vehicle currently entering the parking lot through the host computer, server and other execution devices, so that the vehicle entering the parking lot can directly drive to the vacant parking space for parking. The real-time information of each parking space in the current parking lot can also be displayed on the screen and other peripheral devices so that the vehicle currently entering the parking lot can be informed in time.
[0107] In step 3.3), the surveillance video of the parking lot is input into the deep learning target detection model for processing to determine the parking location of each vehicle.
[0108] Exemplarily, the step of tracking the vehicle through the parking lot video and monitoring the parking position of the vehicle can also be achieved by the following steps, including:
[0109] Step 3.3.1), based on the surveillance video of the parking lot, obtain the video frame at each moment.
[0110] The cameras installed in the parking lot can provide real-time video streams to obtain video frames at every moment.
[0111] In step 3.3.2), the video frames are input into the deep learning target detection YOLOv11 model to extract the target features of each vehicle at each moment.
[0112] Here, the collected parking lot video is input into the trained deep learning target detection model, and the vehicle information is detected by the deep learning target detection model, and the vehicle is located and tracked, and finally the parking position of the vehicle is determined. For the complex scene changes in the parking lot, the YOLOv11 algorithm can be used. YOLOv11 can achieve more accurate vehicle tracking tasks and provide faster processing speed while maintaining the best balance between accuracy and performance. The backbone network of the deep learning target detection YOLOv11 model can be used to extract target features for each frame of video. The target features include the vehicle's license plate features and the parking space features of the parking lot (such as parking space coding), and then the vehicle position is detected based on the target features.
[0113] In step 3.3.3), the target features at all adjacent moments are compared for similarity to obtain the driving trajectory of each vehicle and update the parking position of each vehicle.
[0114] Here, the target feature detection results of each two adjacent frames are compared for similarity to determine whether the target features belong to the same vehicle; the target features belonging to the same vehicle in different video frames are associated, and finally, based on the association results, the driving trajectory of each vehicle is generated and the latest parking position of each vehicle is updated.
[0115] Step 3.4), bind the license plate information and parking location of each vehicle.
[0116] Here, the license plate information and parking location of each vehicle are bound to achieve real-time monitoring and accurate statistics of the parking lot, which is convenient for the parking lot to statistically analyze information such as the number of vacant parking spaces and the number of vehicles entering and leaving the parking lot. For example, when the license plate image is captured and recognized by the camera at the entrance of the parking lot, it is regarded as the number of vehicles entering the parking lot + 1, and the time when the vehicle with this license plate enters the parking lot is recorded; when the license plate image is captured and recognized by the camera at the exit of the parking lot, it is regarded as the number of vehicles leaving the parking lot + 1, and the time when the vehicle with this license plate leaves the parking lot is recorded.
[0117] It should be noted that the deep learning algorithm models used in the embodiments of the present invention are all trained models. The data sets required for training can be composed of public data sets and historical data actually collected. Pre-training is first performed using the public data sets, and then the model parameters are fine-tuned using historical data. All training processes can be performed on a server equipped with an NVIDIA RTX4080 graphics card, and the deep learning networks are all implemented based on the Pytorch framework.
[0118] The embodiments of the present invention can alleviate the technical problems of false detection and missed detection of parking spaces and license plates caused by factors such as the external environment, parking space obstruction, and license plate obstruction in the prior art, and realize intelligent management of parking lots.
[0119] In some embodiments, Figure 5 As shown, an embodiment of the present invention provides a parking lot vehicle management device based on vision, comprising:
[0120] A first processing module inputs a parking space image of each parking space in the parking lot into a deep learning image classification model for processing, and outputs a parking space state probability of each parking space; wherein the deep learning image classification model includes a residual block structure for generating a multi-scale feature map;
[0121] A second processing module performs weighted processing based on the parking space state probability and the parking space state information of each parking space obtained from the parking space sensor to determine the real-time state of each parking space;
[0122] The parking management module monitors the vehicles entering and exiting the parking lot, and feeds back the real-time status of each parking space to the vehicles entering the parking lot, so that the vehicles can park in a vacant parking space based on the real-time status of each parking space.
[0123] Furthermore, the first processing module is used to obtain a parking space image of each parking space in the parking lot, and input the parking space image into a deep learning image classification model; wherein the deep learning image classification model includes a first-stage module, a second-stage module, a third-stage module, a fourth-stage module, a fifth-stage module, a global average pooling layer and a fully connected layer connected in sequence; the first-stage module includes a 7×7 convolutional layer and a 3×3 maximum pooling layer; the second-stage module is composed of three residual block structures, the third-stage module is composed of four residual block structures, and the fourth-stage module is composed of six residual block structures; based on the processing of the residual block structure in the deep learning image classification model, a multi-scale feature map of each parking space is obtained; each of the multi-scale feature maps is processed by the remaining structures in the deep learning image classification model, and the parking space state probability of each parking space is output.
[0124] Furthermore, the first processing module is used to process the feature map of each residual block structure input by first a 1×1 convolution layer, and then process it in parallel by a 1×1 convolution layer, a 3×3 convolution layer and a 5×5 convolution layer to obtain feature maps of three scales; the feature maps of the three scales are spliced in the channel dimension, and then input into the 1×1 convolution layer for compression to obtain a multi-scale feature map of each parking space.
[0125] Furthermore, the second processing module is used to obtain parking space status information of each parking space from a parking space sensor installed at each parking space in the parking lot; wherein the parking space status information is used to characterize whether the parking space detected by the parking space sensor is a vacant parking space or a non-vacant parking space; the parking space status probability of each parking space and the parking space status information are weightedly calculated according to preset weights to determine the real-time status of each parking space; wherein the parking space status probability includes the probability of the parking space being a vacant parking space and the probability of the parking space being a non-vacant parking space.
[0126] Furthermore, the parking management module is used to identify the license plate information of each vehicle entering and exiting the parking lot, and record the entry time and exit time of each vehicle; feedback the real-time status of each parking space to each vehicle entering the parking lot; input the surveillance video of the parking lot into the deep learning target detection model for processing to determine the parking position of each vehicle; and bind the license plate information of each vehicle to the parking position.
[0127] Furthermore, the parking management module is used to obtain the license plate image of each vehicle entering and exiting the parking lot from the entrance and exit cameras of the parking lot; input the license plate image of each vehicle into the UNet image segmentation model, and output the first feature map, second feature map, third feature map, fourth feature map and fifth feature map of each vehicle through the downsampling branch processing of the UNet image segmentation model; process the first feature map through the first feature extraction RFB module to obtain a first effective feature map; process the second feature map through the second feature extraction RFB module to obtain a second effective feature map; process the third feature map through the third feature extraction RFB module to obtain a third effective feature map; process the fourth feature map through the fourth feature extraction RFB module to obtain a fourth effective feature map; and process the fifth feature map through the fifth feature extraction RFB module to obtain a The fifth valid feature map; the first valid feature map is processed by the maximum pooling layer, and then channel-joined with the second valid feature map to obtain a first joint feature map; the first joint feature map is processed by the maximum pooling layer, and then channel-joined with the third valid feature map to obtain a second joint feature map; the second joint feature map is processed by the maximum pooling layer, and then channel-joined with the fourth valid feature map to obtain a third joint feature map; the third joint feature map is processed by the maximum pooling layer, and then channel-joined with the fifth valid feature map to obtain a fourth joint feature map; the fourth joint feature map is compressed by a 1×1 convolution layer to output a final feature map of each of the vehicles; the final feature map of each of the vehicles is input into the channel attention module for processing, and then optical character recognition is performed to obtain the license plate information of each of the vehicles.
[0128] Furthermore, the parking management module is used to obtain video frames at each moment based on the surveillance video of the parking lot; input the video frames into the deep learning target detection YOLOv11 model to extract the target features of each vehicle at each moment; compare the target features of all adjacent moments for similarity to obtain the driving trajectory of each vehicle, and update the parking position of each vehicle.
[0129] An embodiment of the present invention provides an electronic device for implementing an electronic device. In this embodiment, the electronic device may be, but is not limited to, a personal computer (PC), a laptop computer, a monitoring device, a server, or other computer device with analysis and processing capabilities.
[0130] As an exemplary embodiment, see Figure 6 The electronic device 110 includes a communication interface 111, a processor 112, a memory 113 and a bus 114. The processor 112, the communication interface 111 and the memory 113 are connected via the bus 114. The memory 113 is used to store a computer program that supports the processor 112 to execute the method. The processor 112 is configured to execute the program stored in the memory 113.
[0131] The machine-readable storage medium mentioned in this article can be any electronic, magnetic, optical or other physical storage device that can contain or store information, such as executable instructions, data, etc. For example, the machine-readable storage medium can be: RAM (Radom Access Memory), volatile memory, non-volatile memory, flash memory, storage drive (such as hard disk drive), any type of storage disk (such as CD, DVD, etc.), or similar storage medium, or a combination thereof.
[0132] The non-volatile medium may be a non-volatile memory, a flash memory, a storage drive (such as a hard drive), any type of storage disk (such as a CD, DVD, etc.), or a similar non-volatile storage medium, or a combination thereof.
[0133] It can be understood that the specific operation methods of each functional module in this embodiment can refer to the detailed description of the corresponding steps in the above method embodiment, and will not be repeated here.
[0134] The computer-readable storage medium provided in the embodiment of the present invention stores a computer program. When the computer program code is executed, the method described in any of the above embodiments can be implemented. For specific implementation, please refer to the method embodiment, which will not be described in detail here.
[0135] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system and device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0136] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0137] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0138] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the aforementioned embodiments, those of ordinary skill in the art should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the aforementioned embodiments within the technical scope disclosed by the present invention, or can easily conceive of changes, or make equivalent replacements for some of the technical features therein. Such modifications, changes or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention.
Claims
1. A parking lot vehicle management method based on vision, characterized in that: include: Inputting a parking space image of each parking space in the parking lot into a deep learning image classification model for processing, and outputting a parking space state probability of each parking space; wherein the deep learning image classification model includes a residual block structure for generating a multi-scale feature map; Performing weighted processing based on the parking space state probability and the parking space state information of each parking space obtained from the parking space sensor to determine the real-time state of each parking space; The vehicles entering and exiting the parking lot are monitored, and the real-time status of each parking space is fed back to the vehicles entering the parking lot, so that the vehicles can park in a vacant parking space based on the real-time status of each parking space.
2. The method according to claim 1, characterized in that The step of inputting the parking space image of each parking space in the parking lot into the deep learning image classification model for processing and outputting the parking space state probability of each parking space comprises: Obtain a parking space image of each parking space in the parking lot, and input the parking space image into a deep learning image classification model; wherein the deep learning image classification model includes a first-stage module, a second-stage module, a third-stage module, a fourth-stage module, a fifth-stage module, a global average pooling layer and a fully connected layer connected in sequence; the first-stage module includes a 7×7 convolutional layer and a 3×3 maximum pooling layer; the second-stage module consists of three residual block structures, the third-stage module consists of four residual block structures, and the fourth-stage module consists of six residual block structures; Based on the processing of the residual block structure in the deep learning image classification model, a multi-scale feature map of each parking space is obtained; Each of the multi-scale feature maps is processed by the remaining structures in the deep learning image classification model to output the parking space state probability of each of the parking spaces.
3. The method according to claim 2, characterized in that The step of obtaining a multi-scale feature map of each parking space based on the processing of the residual block structure in the deep learning image classification model comprises: The feature map of each residual block structure is first processed by a 1×1 convolution layer, and then processed in parallel by a 1×1 convolution layer, a 3×3 convolution layer, and a 5×5 convolution layer to obtain feature maps of three scales; The feature maps of the three scales are concatenated in the channel dimension, and then input into a 1×1 convolutional layer for compression to obtain a multi-scale feature map of each parking space.
4. The method according to claim 1, characterized in that The step of performing weighted processing based on the parking space state probability and the parking space state information of each parking space obtained from the parking space sensor to determine the real-time state of each parking space includes: Acquiring parking space status information of each parking space from a parking space sensor provided at each parking space in the parking lot; wherein the parking space status information is used to indicate whether the parking space detected by the parking space sensor is a vacant parking space or a non-vacant parking space; The parking space status probability and the parking space status information of each parking space are weightedly calculated according to preset weights to determine the real-time status of each parking space; wherein the parking space status probability includes the probability that the parking space is a vacant parking space and the probability that the parking space is a non-vacant parking space.
5. The method according to claim 1, characterized in that The step of monitoring vehicles entering and exiting the parking lot and feeding back the real-time status of each parking space to the vehicles entering the parking lot comprises: Identify the license plate information of each vehicle entering and leaving the parking lot, and record the entry time and exit time of each vehicle; Feedback the real-time status of each parking space to each vehicle entering the parking lot; Inputting the surveillance video of the parking lot into the deep learning target detection model for processing to determine the parking position of each vehicle; The license plate information of each vehicle is bound to the parking location.
6. The method according to claim 5, characterized in that The step of identifying the license plate information of each vehicle entering and leaving the parking lot comprises: Acquire the license plate image of each vehicle entering and exiting the parking lot from the entrance and exit cameras of the parking lot; Inputting the license plate image of each of the vehicles into a UNet image segmentation model, and outputting a first feature map, a second feature map, a third feature map, a fourth feature map, and a fifth feature map of each of the vehicles through downsampling branch processing of the UNet image segmentation model; The first feature map is processed by a first feature extraction RFB module to obtain a first effective feature map; the second feature map is processed by a second feature extraction RFB module to obtain a second effective feature map; the third feature map is processed by a third feature extraction RFB module to obtain a third effective feature map; the fourth feature map is processed by a fourth feature extraction RFB module to obtain a fourth effective feature map; the fifth feature map is processed by a fifth feature extraction RFB module to obtain a fifth effective feature map; Performing maximum pooling layer processing on the first effective feature map, and then performing channel splicing on the second effective feature map to obtain a first spliced feature map; Performing maximum pooling layer processing on the first spliced feature map, and then performing channel splicing with the third effective feature map to obtain a second spliced feature map; Performing maximum pooling layer processing on the second spliced feature map, and then performing channel splicing with the fourth effective feature map to obtain a third spliced feature map; Performing maximum pooling layer processing on the third spliced feature map, and then performing channel splicing with the fifth effective feature map to obtain a fourth spliced feature map; Compressing the fourth concatenated feature map through a 1×1 convolutional layer to output a final feature map of each vehicle; The final feature map of each vehicle is input into the channel attention module for processing, and then optical character recognition is performed to obtain the license plate information of each vehicle.
7. The method according to claim 5, characterized in that The step of inputting the monitoring video of the parking lot into the deep learning target detection model for processing to determine the parking position of each vehicle includes: Based on the surveillance video of the parking lot, obtaining a video frame at each moment; Input the video frame into the deep learning target detection YOLOv11 model to extract the target features of each vehicle at each moment; The target features at all adjacent moments are compared for similarity to obtain the driving trajectory of each vehicle, and the parking position of each vehicle is updated.
8. A parking lot vehicle management device based on vision, characterized in that: include: A first processing module inputs a parking space image of each parking space in the parking lot into a deep learning image classification model for processing, and outputs a parking space state probability of each parking space; wherein the deep learning image classification model includes a residual block structure for generating a multi-scale feature map; A second processing module performs weighted processing based on the parking space state probability and the parking space state information of each parking space obtained from the parking space sensor to determine the real-time state of each parking space; The parking management module monitors the vehicles entering and exiting the parking lot, and feeds back the real-time status of each parking space to the vehicles entering the parking lot, so that the vehicles can park in a vacant parking space based on the real-time status of each parking space.
9. An electronic device, characterized in that: The method comprises a memory, a processor, and a program stored in the memory and capable of being run on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program.
10. A computer-readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed, the method according to any one of claims 1 to 7 is implemented.