Lightweight parking space detection method and system integrating spatiotemporal consistency and self-supervised learning
By integrating spatiotemporal consistency and self-supervised learning into a lightweight parking space detection method, the robustness and efficiency issues of parking space detection in dynamic environments are solved, and high-precision, low-complexity parking space status recognition is achieved to adapt to complex environmental changes.
Patent Information
- Application Number
- CN202510940558.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing parking space detection methods lack robustness in dynamic environments, are difficult to run efficiently on resource-limited edge devices, and suffer from insufficient detection accuracy.
A lightweight parking space detection method that integrates spatiotemporal consistency and self-supervised learning is adopted. By pre-defining the region of interest, a Gaussian mixture model and frame difference method are combined to generate a foreground mask. A lightweight classification model is used to analyze the parking space status. The model is optimized through self-supervised learning and knowledge distillation, and the background model is dynamically updated to adapt to environmental changes.
The accuracy and adaptability of parking space detection are improved, the computational complexity and storage requirements are reduced, real-time operation on edge devices is ensured, false detections and missed detections are reduced, and the robustness of the system is enhanced.
Smart Images

Figure CN120431553B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of parking lot space detection technology, and in particular to a lightweight parking space detection method and system integrating spatiotemporal consistency and self-supervised learning. Background Art
[0002] With the development of intelligent parking technology, computer vision-based parking space detection has become a core technology in automated parking systems. Traditional parking space detection methods, such as background subtraction and frame subtraction, primarily identify foreground areas by performing pixel-level difference calculations or background modeling on video frames. These methods perform well in static scenes, but are sensitive to factors such as lighting changes, dynamic objects, and obstructions in complex environments, making them prone to detection errors.
[0003] Furthermore, while deep learning-based parking space detection methods have achieved significant progress in accuracy, they are computationally intensive and require numerous parameters, making them difficult to run efficiently on resource-constrained edge devices. To address this issue, researchers have proposed lightweight deep learning models that compress models through techniques such as pruning, quantization, or knowledge distillation to reduce computational overhead. While these methods improve inference efficiency, their accuracy in detecting parking space status changes in dynamic environments remains to be improved.
[0004] Existing parking space detection methods generally lack robustness against background variations and fail to fully leverage the advantages of spatiotemporal consistency and self-supervised learning. Therefore, ensuring detection accuracy while improving the real-time and adaptability of the model has become a pressing technical challenge in the field of parking space detection. Summary of the Invention
[0005] The problem to be solved by the present invention is to provide a lightweight parking space detection method and system that integrates spatiotemporal consistency and self-supervised learning, accurately selects the area to be detected in a dynamic environment, and makes adaptive adjustments according to real-time environmental changes to improve the parking space detection accuracy.
[0006] The present invention adopts the following technical solution: a lightweight parking space detection method integrating spatiotemporal consistency and self-supervised learning, comprising the following steps:
[0007] Step 1: Capture parking lot video streams using cameras, predefine regions of interest (ROIs) for each parking space in the parking lot, capture images of each ROI, and perform initial state classification on all ROIs using a lightweight classification model.
[0008] Step 2: Based on the image of the area of interest of each parking space, a Gaussian mixture model (GMM) is used to build a background model, which is compared with the current frame image to extract the first foreground mask;
[0009] Step 3: Calculate the pixel difference between the current frame image and the previous frame image by using the frame difference method to obtain the second foreground mask;
[0010] Step 4: Fuse the first foreground mask and the second foreground mask, remove noise through morphological processing, and generate the final foreground mask;
[0011] Step 5: Select the area to be detected based on spatiotemporal consistency, calculate the image change ratio of each ROI, and determine whether a state change occurs;
[0012] Step 6: Optimize and pre-train the lightweight classification model, and call the pre-trained lightweight classification model to perform in-depth classification analysis of parking space status on the detected ROI with state changes;
[0013] Step 7: Dynamically adjust the background model based on the depth analysis results, and update the background of each parking space in real time to adapt to lighting changes and environmental dynamics.
[0014] Preferably, in step 1, based on the camera position in the parking lot, the area of interest of each parking space is predefined ;
[0015] The lightweight classification model is MobileNetV3, and the lightweight classification model is used to perform initial state classification on the areas of interest of all parking spaces to obtain a preliminary classification result: the parking space is free or occupied.
[0016] Preferably, in step 2, a background mask is established using a Gaussian mixture model to obtain Background model at this moment , compare the current frame With background model , generate the first foreground mask .
[0017] Preferably, in step 3, the current frame is calculated by the frame difference method With the previous frame Pixel difference , according to the difference image , generate the second foreground mask .
[0018] Preferably, in step 4, the first foreground mask and the second foreground mask are fused by a logical AND operation to obtain a fused foreground mask , for the fused foreground mask Perform morphological processing to remove noise and fill holes to obtain the final foreground mask .
[0019] Preferably, in step 5, the area to be detected is selected based on spatiotemporal consistency, and the sub-steps include:
[0020] Step 5.1: Calculate the image change ratio of each parking space region of interest, and calculate the image change ratio of each parking space based on the pixel change of the final foreground mask. , calculate the change ratio ;
[0021] Step 5.2: Use the set change threshold , classify the state of the area of interest of each parking space, determine whether the parking space state has changed, and perform in-depth classification analysis when the parking space state changes.
[0022] Preferably, in step 6, the lightweight classification model is optimized, knowledge is acquired from the teacher model through distillation learning, the student model is guided to be optimized by calculating the distillation loss function, and the lightweight classification model is pre-trained through self-supervised learning. The tasks include but are not limited to autoencoder training and feature consistency loss optimization. The pre-trained lightweight classification model is used to perform in-depth analysis of parking space status classification.
[0023] Preferably, in step 7, the background model is dynamically adjusted according to the depth analysis results. The background model is updated using an online learning mechanism. Based on the detected foreground information, the foreground changes of consecutive frames are verified through spatiotemporal consistency analysis. If the background change is greater than the threshold, the background model is adjusted to adapt to the environmental changes to ensure the accuracy of background modeling.
[0024] The technical solution of the present invention also provides: a lightweight parking space detection system integrating spatiotemporal consistency and self-supervised learning, for implementing any of the above methods, comprising: a perception module, a foreground extraction module, a region selection module, a lightweight classification module, a background update module, and a distillation training module;
[0025] The perception module is used to collect parking lot video streams through cameras;
[0026] The foreground extraction module is used to establish a background model, compare it with the current frame image, and extract a first foreground mask; calculate the pixel difference between the current frame image and the previous frame image by a frame difference method to obtain a second foreground mask;
[0027] The region selection module is used to select the region to be detected based on spatiotemporal consistency;
[0028] The lightweight classification module is used to classify the parking space status of the area to be detected;
[0029] The background update module is used to dynamically update the background model according to the depth analysis result;
[0030] The distillation training module is used to optimize the lightweight classification model through self-supervised knowledge distillation.
[0031] The technical solution of the present invention further provides: an electronic device, comprising:
[0032] one or more processors;
[0033] a storage device having one or more programs stored thereon;
[0034] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above-mentioned lightweight parking space detection methods that integrate spatiotemporal consistency and self-supervised learning.
[0035] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:
[0036] 1. The lightweight parking space detection method of the present invention captures the time series information of parking space status changes through spatiotemporal consistency analysis. It can more accurately select the area to be detected in a dynamic environment, effectively reduce false detections and missed detections caused by factors such as lighting changes, background dynamics, and parking space obstructions, and improve the accurate judgment of parking space occupancy status.
[0037] 2. The lightweight parking space detection method of the present invention significantly reduces the computational complexity and storage requirements of the model by combining self-supervised learning and knowledge distillation with a lightweight model. At the same time, it improves the generalization ability of the model through self-supervised pre-training without relying on a large amount of labeled data. And through distillation technology, the knowledge of the teacher model is effectively transferred to the student model, so that the detection accuracy is retained while the inference speed is greatly improved, ensuring that it can run in real time on resource-limited edge devices.
[0038] 3. The lightweight parking space detection method of the present invention adopts an online learning mechanism to dynamically update the background model. Unlike the traditional static background modeling method, the background model can be adaptively adjusted according to real-time environmental changes to maintain the accuracy of background modeling, thereby improving the robustness and adaptability of the system in complex practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 This is a flowchart of the lightweight parking space detection method of the present invention;
[0040] Figure 2 A flow chart for determining whether a background model needs to be updated is established for the present invention;
[0041] Figure 3 This is the knowledge distillation structure diagram of the parking space classification model of the present invention. DETAILED DESCRIPTION
[0042] In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions of the application are further elaborated in detail below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments involved in the present invention. All non-innovative embodiments of other researchers in this field on this embodiment fall within the scope of protection of the present invention. At the same time, the step numbers in the embodiments of the present invention are only set for the convenience of explanation and description, and the order between the steps is not limited in any way. The execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0043] Example 1
[0044] This embodiment provides a lightweight parking space detection method based on spatiotemporal consistency and self-supervised knowledge distillation. Figure 1 As shown, the following steps are included:
[0045] Step 1: Obtain a continuous video stream from a camera fixed at the entrance and exit of the parking lot, with a frame rate of 15fps and a resolution of 1920×1080. Predefine the region of interest (ROI) for each parking space in the parking lot based on the fixed camera position, and perform an initial state classification on all ROIs using a lightweight classification model.
[0046] Step 2: Use the Gaussian mixture model to build a background model, compare it with the current frame image, and extract the first foreground mask;
[0047] Step 3: Calculate the pixel difference between the current frame and the previous frame using the frame difference method to obtain the second foreground mask;
[0048] Step 4: Fusing the first and second foreground masks obtained in steps 2 and 3, removing noise through morphological processing, and generating a final foreground mask;
[0049] Step 5: Select the area to be detected based on spatiotemporal consistency, calculate the change ratio of each ROI area, and determine whether its state changes;
[0050] Step 6: Optimize and pre-train the lightweight classification model. For the ROI where changes are detected, call the pre-trained lightweight classification model to perform in-depth classification analysis and detection of parking space status.
[0051] Step 7: Dynamically adjust the background model based on the detection results and update the background in real time to adapt to lighting changes and environmental dynamics.
[0052] As a specific implementation of this embodiment, in step 1, the lightweight classification model is MobileNetV3, and the region of interest of each parking space is predefined, which is expressed as .
[0053] As a specific implementation of this embodiment, step 2 specifically includes the following steps:
[0054] Step 2.1: Use the Gaussian mixture model to create a background mask. The update formula is as follows:
[0055] ;
[0056] in, is the learning rate, is the current frame, Respectively and Background model of the moment.
[0057] Step 2.2, by comparing the current frame With background model , generate the first foreground mask :
[0058] ;
[0059] in, is the threshold value used to determine whether the pixel belongs to the foreground. Indicates the image coordinates.
[0060] As a specific implementation of this embodiment, step 3, such as Figure 2 As shown, the specific steps include:
[0061] Step 3.1: Calculate the current frame by frame difference method With the previous frame Pixel difference
[0062] ;
[0063] Step 3.2: Based on the difference image , generate frame difference foreground mask
[0064] ;
[0065] in, is the frame difference threshold, which is used to determine whether a pixel has changed significantly.
[0066] As a specific implementation of this embodiment, step 4 specifically includes the following steps:
[0067] Step 4.1: Fuse the foreground mask using logical AND operation
[0068] ;
[0069] in, Represents a logical AND operation.
[0070] Step 4.2: Mask the fused foreground Perform morphological processing (such as dilation and erosion) to remove noise and fill holes to obtain the processed foreground mask :
[0071] ;
[0072] in, Represents opening-closing morphological filtering.
[0073] As a specific implementation of this embodiment, step 5 specifically includes the following steps:
[0074] Step 5.1: The calculation of the change ratio is based on the pixel change of the foreground mask, and the set threshold is used to classify the state of the ROI to determine whether the parking space needs to be subjected to depth classification analysis.
[0075] Step 5.2: For each parking space , calculate the change ratio :
[0076] ;
[0077] in, is the total number of pixels within the ROI, is the final foreground mask.
[0078] Step 5.3: Set the change threshold , determine whether the parking space status has changed:
[0079] ;
[0080] in, Is the result of whether the parking space status has changed, when When it is 1, deep classification analysis is required.
[0081] As a specific implementation of this embodiment, step 6 specifically includes the following steps:
[0082] Step 6.1: Pre-train the lightweight classification model through self-supervised learning. Pre-training tasks include but are not limited to autoencoder training and feature consistency loss optimization:
[0083] ;
[0084] in, and is the feature representation of different views of the same parking space, is a similarity function (such as cosine similarity), is the temperature parameter, is the batch size.
[0085] Furthermore, the pre-trained lightweight classification model is used to perform model forward reasoning and output the parking space occupied-vacant classification results.
[0086] Step 6.2, such as Figure 3 As shown in Figure 2, knowledge is acquired from a high-performance teacher model (such as ResNet-10) through distillation learning. The distillation process involves calculating a distillation loss function to guide the optimization of the student model:
[0087] ;
[0088] in, 、 Current frame Teacher model and student model; is the cross entropy loss, which is used to measure the difference between the student model's prediction and the true label the differences between; is the Kullback-Leibler divergence, which is used to measure the difference in the output distribution of the teacher model and the student model; and is the weight coefficient, is the temperature parameter.
[0089] Preferably, in this example , balance the two losses; take , used to soften the probability distribution and improve the distillation effect.
[0090] As a specific implementation of this embodiment, in step 7, the background model is updated using an online learning mechanism. If the background change is greater than a threshold value based on the detected foreground information, the background model is adjusted to adapt to the environmental changes and ensure the accuracy of the background modeling:
[0091] ;
[0092] in, 、 Respectively and Background model of the moment.
[0093] Example 2
[0094] Based on the same inventive concept as the first embodiment, this embodiment provides a lightweight parking space detection system that integrates spatiotemporal consistency and self-supervised learning, including: a perception module, a foreground extraction module, a region selection module, a lightweight classification module, a background update module, and a distillation training module;
[0095] Among them, the perception module is used to collect parking lot video streams through cameras;
[0096] The foreground extraction module is used to establish a background model, compare it with the current frame image, and extract a first foreground mask; calculate the pixel difference between the current frame image and the previous frame image through the frame difference method to obtain a second foreground mask;
[0097] A region selection module is used to select the region to be detected based on spatiotemporal consistency;
[0098] A lightweight classification module is used to classify the parking space status of the area to be inspected;
[0099] Background update module, used to dynamically update the background model according to the depth analysis results;
[0100] A distillation training module for optimizing lightweight classification models through self-supervised knowledge distillation.
[0101] Example 3
[0102] Based on the same inventive concept as other embodiments, this embodiment provides a computer device, including: a memory for storing a computer program; a processor for executing the computer program to implement any of the steps of the above-mentioned lightweight parking space detection method that integrates spatiotemporal consistency and self-supervised learning.
[0103] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0104] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0105] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0106] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0107] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A lightweight parking space detection method that integrates spatiotemporal consistency and self-supervised learning, characterized by: The steps include: Step 1: Capture parking lot video streams using cameras, predefine regions of interest for each parking space in the parking lot, capture images of the regions of interest for each parking space, and perform initial state classification on the regions of interest for all parking spaces using a lightweight classification model. Step 2: Based on the image of the area of interest of each parking space, a background model is established using a Gaussian mixture model, which is compared with the current frame image to extract the first foreground mask; Step 3: Calculate the pixel difference between the current frame image and the previous frame image by using the frame difference method to obtain the second foreground mask; Step 4: Fuse the first foreground mask and the second foreground mask, remove noise through morphological processing, and generate the final foreground mask; Step 5: Select the area to be detected based on spatiotemporal consistency, calculate the image change ratio of the area of interest of each parking space, and determine whether a state change has occurred. This includes the following sub-steps: Step 5.1: Calculate the image change ratio of each parking space region of interest, and calculate the image change ratio of each parking space based on the pixel change of the final foreground mask. , calculate the change ratio : ; in, is the total number of pixels in the parking space area of interest, is the final foreground mask obtained in step 4; Step 5.2: Use the set change threshold , classify the state of the area of interest of each parking space and determine whether the state of the parking space has changed: ; in, Is the result of whether the parking space status has changed, when When it is 1, deep classification analysis is performed; Step 6: Optimize and pre-train the lightweight classification model, and call the pre-trained lightweight classification model to perform in-depth classification analysis of parking space status on the detected parking space interest area with status changes; Step 7: Dynamically adjust the background model based on the results of the deep classification analysis of the parking space status, and update the background of each parking space in real time to adapt to lighting changes and environmental dynamics.
2. The lightweight parking space detection method integrating spatiotemporal consistency and self-supervised learning according to claim 1 is characterized in that: In step 1, based on the camera positions in the parking lot, pre-define the area of interest for each parking space ; The lightweight classification model is MobileNetV3, and the lightweight classification model is used to perform initial state classification on the areas of interest of all parking spaces, and the parking space classification results are obtained as free or occupied.
3. The lightweight parking space detection method integrating spatiotemporal consistency and self-supervised learning according to claim 2 is characterized in that: In step 2, the background model is established using the Gaussian mixture model. The formula is as follows: ; in, is the learning rate, for The current frame at the moment, Respectively and The background model of the moment; Compare current frame With background model , generate the first foreground mask , the formula is as follows: ; in, is the threshold value used to determine whether the pixel belongs to the foreground. Indicates the image coordinates.
4. The lightweight parking space detection method integrating spatiotemporal consistency and self-supervised learning according to claim 3 is characterized in that: In step 3, the current frame is calculated by the frame difference method With the previous frame Pixel difference , the formula is as follows: ; Based on the difference image , generate the second foreground mask : ; in, is the frame difference threshold, which is used to determine whether a pixel has changed significantly.
5. The lightweight parking space detection method integrating spatiotemporal consistency and self-supervised learning according to claim 4 is characterized in that: In step 4, the first foreground mask and the second foreground mask are fused by logical AND operation to obtain the fused foreground mask : ; in, Represents logical AND operation; Foreground mask after fusion Perform morphological processing to remove noise and fill holes to obtain the final foreground mask : ; in, Represents opening-closing morphological filtering.
6. The lightweight parking space detection method integrating spatiotemporal consistency and self-supervised learning according to claim 3 is characterized in that: In step 6, the lightweight classification model is optimized by acquiring knowledge from the teacher model through distillation learning. The distillation loss function is calculated to guide the optimization of the student model, which is expressed as follows: ; in, is the cross entropy loss, which is used to measure the difference between the student model's prediction and the true label the differences between; is the Kullback-Leibler divergence, which is used to measure the difference in the output distribution of the teacher model and the student model; and is the weight coefficient, balancing the two parts of loss; is the temperature parameter used to soften the probability distribution; 、 Current frame Teacher model and student model; The lightweight classification model is pre-trained through self-supervised learning. The tasks include but are not limited to autoencoder training and feature consistency loss optimization, which are expressed as follows: ; in, and is the feature representation of different views of the same parking space, is the similarity function, is the temperature parameter, is the batch size; Through the pre-trained lightweight classification model, the model forward reasoning is performed and the parking space occupied or idle classification results are output.
7. The lightweight parking space detection method integrating spatiotemporal consistency and self-supervised learning according to claim 3 is characterized in that: In step 7, the background model is dynamically adjusted based on the depth analysis results. The background model is updated using an online learning mechanism. Based on the detected foreground information, the foreground changes of consecutive frames are verified through spatiotemporal consistency analysis. If the background change is greater than the threshold, the background model is adjusted: ; in, 、 Respectively and Background model of the moment.
8. A lightweight parking space detection system integrating spatiotemporal consistency and self-supervised learning, for implementing the method according to any one of claims 1 to 7, characterized in that: include: Perception module, foreground extraction module, region selection module, lightweight classification module, background update module, and distillation training module; The perception module collects parking lot video streams through cameras, predefines regions of interest for each parking space in the parking lot, collects images of the regions of interest for each parking space, and performs initial state classification on the regions of interest for all parking spaces using a lightweight classification model; The foreground extraction module uses a Gaussian mixture model to build a background model based on the image of the area of interest of each parking space, compares it with the current frame image, and extracts a first foreground mask; and calculates the pixel difference between the current frame image and the previous frame image using a frame difference method to obtain a second foreground mask; The first foreground mask and the second foreground mask are fused, noise is removed by morphological processing, and a final foreground mask is generated; The region selection module is used to select the area to be detected based on spatiotemporal consistency, calculate the image change ratio of each parking space area of interest, and determine whether a state change occurs: Calculate the image change ratio of each parking space area of interest, and based on the pixel change of the final foreground mask, , calculate the change ratio ; Use the set change threshold to classify the state of the area of interest of each parking space, determine whether the parking space state has changed, and perform in-depth classification analysis when the result is 1; The lightweight classification module is used to classify the parking space status of the detected area and call the pre-trained lightweight classification model to perform in-depth classification analysis of the parking space status for the parking space interest area detected to have a state change; The background update module is used to dynamically adjust the background model according to the results of the deep classification analysis of the parking space status, and update the background of each parking space in real time to adapt to lighting changes and environmental dynamics; The distillation training module is used to optimize and pre-train the lightweight classification model through self-supervised knowledge distillation.
9. An electronic device, characterized in that: include: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the lightweight parking space detection method that integrates spatiotemporal consistency and self-supervised learning as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Dynamic adaptive depth camera occlusion detection method, system and device, and storage medium
CN120182813A
Adaptive modeling method for background image, detecting method and system for illegal-stopping and parking vehicle using it
KR101038650B1