Parking space reasoning model in complex scene

By combining stacked hourglass networks with RESA modules and local-global distillation algorithms, along with corner coordinate algorithms, the problems of missed detection and false detection in parking space detection under complex scenarios are solved, improving the accuracy of parking space detection and the reliability of automatic parking systems.

CN121053408APending Publication Date: 2025-12-02ZHONGKE HUIYAN (TIANJIN) ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511151038.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

Existing parking space detection methods are prone to missed detections and false detections in complex scenarios, especially when lighting conditions are complex or the parking space is partially obscured, which leads to a decrease in the accuracy and reliability of automatic parking systems.

Method used

A stacked hourglass network and a RESA module are used to process images. A lightweight student hourglass module is trained using a local-global distillation algorithm. The corner coordinate algorithm is used to match key parking space information, thereby reducing the false negative and false positive rates.

Benefits of technology

It effectively improves the accuracy and robustness of parking space detection, especially under complex lighting conditions and occlusion scenarios, significantly reducing the missed detection rate and false detection rate, thereby improving the practicality and user experience of the automatic parking system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121053408A_ABST
    Figure CN121053408A_ABST
Patent Text Reader

Abstract

The invention provides a parking space reasoning model in a complex scene, which is characterized by comprising the following steps of: 1, processing an input image by using a stacked hourglass network (stacked hourglass network) and RESA (Recurrent Feature Shift Aggregate) module, and extracting context features; step 2, matching key information of the parking space by applying an algorithm based on angular point coordinates, wherein the key information comprises a position, a type and an occupancy state; step 3, training a lightweight student hourglass module by using a local-global distillation (FGD) algorithm, reducing parameter quantity and maintaining generalization ability and detection precision; according to the invention, the limitation of the existing parking space detection technology is analyzed, and the system is comprehensively and meticulously optimized from multiple key dimensions. Therefore, the missing detection rate and the false detection rate are effectively reduced, and the practicability and the user experience of the whole automatic parking system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of parking space detection methods, and in particular to a parking space reasoning model for complex scenarios. Background Technology

[0002] With the deep integration of artificial intelligence, sensor technology, and control theory, integrated driving assistance systems are rapidly moving from the laboratory to real-world roads, profoundly reshaping human transportation patterns. As a key component of integrated driving assistance systems, automated parking systems achieve autonomous parking without human intervention through multi-sensor data fusion, intelligent decision-making and planning, and precise control execution, significantly improving driving safety and convenience. In parking space reasoning, parking space detection is a core component, and its accuracy and robustness directly determine the overall performance of the system.

[0003] Traditional parking space detection methods mostly rely on ultrasonic or infrared sensors to determine parking space availability by collecting distance information around the vehicle. However, these methods have many limitations. For example, ultrasonic sensors are prone to significant errors at vehicle corners, and the signal propagation during measurement is easily affected by environmental interference, resulting in inaccurate measurement results. They struggle to clearly capture detailed information about the parking space and cannot provide comprehensive and accurate parking environment data for automatic parking systems.

[0004] The rise of vision-based deep learning methods has revolutionized the field of parking space detection. By constructing multi-layered neural networks, deep learning can learn complex feature patterns from massive amounts of data, achieving significant results in computer vision tasks such as object detection and image segmentation. Through learning from large amounts of parking space image data, automated parking space detection can be achieved. Compared to traditional methods, deep learning has significant advantages in accuracy, adaptability, and efficiency. It can learn more complex parking space features, improving detection accuracy; it can adapt to the parking space detection needs of different scenarios, possessing good generalization ability; and it can achieve automated detection, effectively reducing costs and time consumption.

[0005] With further technological advancements, the Around View Monitor (AVM) system, as a key technology, plays an increasingly important role in automated parking systems. The AVM system acquires multi-angle images using fisheye cameras around the vehicle and generates a 360° virtual panoramic view of the vehicle's surroundings using image stitching and distortion correction techniques. This provides high-resolution, wide-view input data for deep learning models. This data fusion method significantly improves the system's ability to perceive complex parking environments, such as narrow parking spaces or scenes obstructed by obstacles, enabling clearer capture of parking space boundaries and obstacle positions. However, despite the achievements of deep learning-based parking space detection methods, many challenges remain in practical applications.

[0006] During parking, vehicles often encounter situations where parking spaces are partially obscured, such as by other vehicles or obstacles in front of or behind the space, resulting in incomplete display of the parking space. This presents a new challenge for parking space detection algorithms. Furthermore, complex lighting conditions are also a significant contributing factor. Issues such as shadows and overexposure can easily lead to incomplete or blurred display of key features like parking space lines, causing the recognition system to malfunction and greatly increasing the difficulty of parking space detection. According to relevant research, in some complex scenarios, the false negative and false positive rates of existing parking space detection methods increase significantly, seriously affecting the reliability and practicality of automatic parking systems. Summary of the Invention

[0007] In view of the above technical problems, the present invention provides a parking space reasoning model for complex scenarios that solves the problems of missed detection and false detection caused by the variability of scenarios in the parking space detection process. Especially in challenging scenarios such as complex lighting conditions and partially obscured parking spaces, it effectively reduces the missed detection rate and false detection rate, and improves the practicality and user experience of the entire automatic parking system.

[0008] This invention provides a parking space reasoning model for complex scenarios, characterized by the following steps:

[0009] Step 1: Process the input image using the stacked hourglass network and the RESA (Recurrent Flour Shift Aggregator) module to extract contextual features;

[0010] Step 2: Apply an algorithm based on corner coordinates to match key parking space information, including location, type, and occupancy status;

[0011] Step 3: Train the lightweight student hourglass module using the local-global distillation (FGD) algorithm to reduce the number of parameters while maintaining generalization ability and detection accuracy.

[0012] The specific steps of step one include:

[0013] 1) Network adjustment: The input image is compressed by the SPD-Conv layer. Through the joint processing of space-to-depth transformation and non-strut convolution, the output feature map has a resolution of 128×128 and 48 channels.

[0014] 2) Cyclic Feature Transformation Aggregator: The RESA module utilizes the strong shape prior of parking spaces and the cross-row and column spatial information of pixels to extract features related to parking spaces. This module cyclically moves the features of the 3D feature map in four directions in sequence, so that each pixel collects global information. Given a 3D feature map tensor X, which contains three dimensions: the number of channels C, the number of rows H, and the number of columns W, the module iterates multiple times to gradually aggregate global information in the vertical and horizontal directions, and finally updates each element of the feature map to enhance the expressive power of parking space features.

[0015] The specific steps of step two are as follows: the output of the RESA module is used as the input for subsequent processing and is introduced into the prediction network. The prediction network is constructed by connecting four hourglass modules in series. The hourglass module is a key component of the network and includes a sampling module, a downsampling module, an upsampling module, and three output branches, which are designed to accurately predict key points on the parking line, parking angle information, and the occupancy status of the parking space.

[0016] The same sampling module uses Conv(1 / 0 / 1) convolution operation; the downsampling module uses Conv(3 / 1 / 2) convolution operation; the upsampling module uses Transposed Conv(3 / 1 / 2) transposed convolution operation; and each module is connected to a Prelu activation function and batch normalization after the convolution operation.

[0017] The three output branches include a confidence branch, an angle branch, and a state branch. Each output branch consists of an SE attention mechanism, a convolutional layer, and a specific activation function.

[0018] The SE attention mechanism consists of two steps: Squeeze and Excitation;

[0019] The squeezing step is as follows: First, the input feature map is compressed into a vector using a global average pooling operation; then, it is mapped to a smaller vector using a fully connected layer.

[0020] The excitation step is as follows: Each element of the vector is compressed to between 0 and 1 using a sigmoid function, and then multiplied with the original input feature map to obtain a weighted feature map. The formulas for the activation functions Sigmoid and Tanh are as follows:

[0021]

[0022]

[0023] Where sigmoid(x) is the Sigmoid function with input variable x; tanh(x) is the Tanh function with input variable x.

[0024] The confidence branch uses a combination of SE-Block+Conv+Sigmoid to output the confidence level of parking space-related information; the angle branch uses a combination of SE-Block+Conv+Tanh to specifically predict parking space orientation information; and the state branch uses a combination of SE-Block+Conv+Sigmoid to determine the occupancy status of the parking space.

[0025] The specific steps of step three are as follows: local distillation separates parking space and background information, guides the student network to focus on the key pixels and channels of the teacher network, and captures the core information in parking space detection; global distillation reconstructs the relationship between different pixels and transmits these relationships from the teacher network to the student network to compensate for the global information that may be missing during the local distillation process.

[0026] The joint training of the local-global distillation (FGD) algorithm includes multiple loss functions, specifically confidence loss, angle loss, occupied state loss, knowledge distillation loss, and total loss.

[0027] The confidence loss uses Focal Loss instead of the standard Cross-Entropy Loss. Focal Loss, by introducing a modulation factor, can adaptively reduce the weight of easily classified samples, making the model pay more attention to difficult-to-classify positive samples. The specific expression is as follows:

[0028] L pos =α pos ∑((v gt -y pred ) 2 log(y pred ))

[0029]

[0030] Among them, L pos and L negy represents the confidence loss for positive and negative samples, respectively. pred y is the network's predicted value. gt The label value is the weight coefficient α of the positive sample. pos =2, the weighting coefficient α of the negative samples neg =0.1;

[0031] The angle loss uses a cosine-based loss function to measure the difference between the predicted angle and the true angle, as shown in the following expression:

[0032]

[0033]

[0034] in, The angles formed by the parking space spacing lines with the positive X-axis and positive Y-axis in the pixel coordinate system represent the angles. Based on the constrained angle loss, the angles of these two features can be accurately predicted, and the orientation of the parking space can be inferred in the post-processing algorithm. pred θ1, cos pred θ2 is the cosine value of the angle predicted by the network, cos gt θ1, cos gt θ2 represents the corresponding label value, and β is the weighting coefficient for the angle loss. θ1 =θ β2 =1;

[0035] The occupancy loss uses the mean squared error (MSE) loss function to measure the difference between the predicted and actual values, and its specific form is as follows:

[0036] L free =γ free ∑(Free gt -Free pred ) 2

[0037] L occupied =γ occupied ∑(occupied gt -Occupied pred ) 2

[0038] Among them, L free and L occupied These represent the losses for the idle and occupied states, respectively. pred Occupied pred It is a network prediction value, Free gt Occupied gt For the corresponding label value, the weight coefficient γ of the state loss is used. free =γoccupied =1, the mean squared error loss function can intuitively reflect the degree of deviation between the predicted occupancy status and the actual occupancy status;

[0039] The knowledge distillation loss is a weighted sum of the local distillation loss and the global distillation loss, and its expression is as follows.

[0040] L distillation =δ focal L focal +δ global L global

[0041] Among them, L distillation For knowledge distillation loss, L focal and L global These represent local distillation loss and global distillation loss, respectively, with the weighting coefficient δ for knowledge distillation loss. focal =δ global =1;

[0042] The total loss function is a weighted sum of confidence loss, angle loss, occupied state loss, and knowledge distillation loss. L2 regularization loss is also introduced to prevent overfitting and enhance the model's generalization ability. Its expression is as follows:

[0043]

[0044] Among them, L total For the total loss, Let be the L2 regularization loss, representing the sum of squares of the parameters in the weight vector. The weight coefficients of the L2 regularization loss are λ = 0.001.

[0045] The specific steps of the corner coordinate algorithm in step two include:

[0046] 1) Distinguish between the detected entry line corner points and exit line corner points, and perform feature matching on the entry line corner points;

[0047] 2) In the feature point matching stage, feature points are paired based on the actual distance X of the entrance line corner point. The actual distance I of the cutoff line is determined using distance X, and the corner point of the parking space cutoff line is calculated.

[0048] 3) When fitting the parking space area, the detected parking space entrance line corner point and the calculated cutoff line corner point are combined to form a parking space; in addition, if the cutoff line corner point detected by the deep learning network model proposed in this invention is close to the calculated cutoff line corner point, the detected parking space entrance line corner point and the two detected cutoff line corner points are directly combined to form a parking space, and on this basis, the parking space occupancy status is detected.

[0049] The beneficial effects of this invention are as follows: This invention provides a parking space reasoning model for complex scenarios, aiming to effectively solve the problems of missed detection and false detection caused by the variability of scenarios during parking space detection. This is particularly true in challenging scenarios such as complex lighting conditions (e.g., shadow occlusion, underexposure, or overexposure) and partially obscured parking spaces (e.g., a vehicle blocking a parking space during parking). These problems significantly increase the difficulty of parking space detection and reduce the accuracy and reliability of the automatic parking system. To achieve this goal, this invention analyzes the limitations of existing parking space detection technologies and comprehensively and meticulously optimizes the system from multiple key dimensions. This effectively reduces the missed detection rate and false detection rate, improving the practicality and user experience of the entire automatic parking system.

[0050] By utilizing the hourglass network and the RESA module, this invention can deeply mine and extract contextual features, enhancing the model's ability to perceive parking space features in complex parking scenarios and effectively improving the accuracy of parking space detection. This method is particularly suitable for handling occluded areas, blurred parking spaces, and blurred images, thereby improving the robustness of parking space detection.

[0051] To meet the multi-dimensional requirements of parking space reasoning in terms of location, type, and occupancy status, this invention proposes a post-processing matching algorithm based on corner coordinates. This algorithm significantly improves the completeness of parking space detection and ensures the accuracy of detection results by accurately identifying various key information of the parking space.

[0052] This invention applies the local-global distillation algorithm to the training of a lightweight student hourglass module. By constructing a feature supervision system, it drives the student model to accurately learn the core knowledge of the teacher model, achieving a significant reduction in the number of model parameters while effectively maintaining its generalization ability and detection accuracy. This technique not only improves the model's efficiency but also ensures its performance. Attached Figure Description

[0053] Figure 1 This is a diagram of the PIPS-Net network architecture of the present invention;

[0054] Figure 2 This is a schematic diagram of the cyclic feature conversion aggregator of the present invention;

[0055] Figure 3 This is a schematic diagram of the hourglass module of the present invention;

[0056] Figure 4 This is a flowchart of the parking space detection process of the present invention;

[0057] Figure 5 This is a schematic diagram illustrating the parking space angle prediction method of the present invention.

[0058] Figure 6This is a schematic diagram illustrating the parking space (vacant / occupied) status prediction of the present invention;

[0059] Figure 7 This is a schematic diagram showing the parking space detection results in different environments according to the present invention. Detailed Implementation

[0060] Example 1

[0061] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings.

[0062] To address the limitations of existing detection methods in complex scenarios, this invention innovatively introduces a Stacked Hourglass Network as the core feature extraction module. Through the feature recursive enhancement mechanism of a multi-level hourglass network, it is deeply integrated with the RESA (Recurrent Feature Shift Aggregator) module to construct a hierarchical feature representation system with enhanced spatial awareness.

[0063] Building upon this, model compression and parameter optimization are achieved through a knowledge distillation strategy. A novel network architecture, PIPS-Net, suitable for parking space detection, is constructed, such as... Figure 1 As shown.

[0064] This invention not only achieves accurate detection of key points in parking spaces, but also completes parking space spatial structure analysis and occupancy status determination through a post-processing mechanism, forming a complete detection solution from feature extraction to decision output, such as... Figure 4 As shown.

[0065] The parking space detection model proposed in this invention adopts a multi-stage processing architecture. In the image preprocessing stage, the input image is first compressed by the SPD-Conv layer to uniformly adjust it to a resolution of 128×128. Subsequently, the Recurrent Feature Shift Aggregator fuses local detail features and global contextual information into the normalized image. These optimized features are input into the backbone network composed of four stacked hourglass modules, and accurate parking space-related information is generated through multi-scale feature fusion and keypoint regression. However, the complexity of this network architecture places high demands on hardware, requiring high-end GPUs to ensure smooth operation. To overcome hardware limitations, this invention adopts an object detection-based FGD distillation method. A student network is constructed using a single hourglass module, and lightweight training is carried out using FGD distillation technology. FGD separates foreground and background features, enabling the student network to learn key information from the teacher network and accurately reconstruct pixel relationships, ensuring improved model lightweighting and application feasibility without sacrificing performance. In the post-processing stage, based on information such as the coordinates of the parking space boundary corners output by the model, a parking space framework is constructed through feature point matching and region fitting. By combining the predicted angles and the relationship between the corners, a complete reconstruction of the parking space spatial structure is achieved. Finally, based on the occupancy information instantiated by the network, the parking space occupancy status is quickly determined.

[0066] 5.1 Network Architecture

[0067] 1) Network Adjustment: To address the network optimization issue, this study employs SPD-Conv for structural improvement, optimizing computational efficiency while maintaining feature extraction capabilities. This module takes a 1024×1024 resolution RGB image as input and outputs a 128×128 resolution feature map with 48 channels through collaborative processing of space-to-depth transformation and non-stretch convolution. Specifically, the SPD layer transforms spatial information to the channel dimension through pixel block rearrangement, effectively preserving the detailed features of the high-resolution image; subsequently, a convolution with a stride of 1 is used for fine feature extraction, avoiding information loss in traditional downsampling.

[0068] 2) Cyclic Feature Transformation Aggregator: The RESA module utilizes the strong shape prior of parking spaces and the cross-row and column spatial information of pixels to more effectively extract features related to parking spaces, thereby improving the detection and recognition capabilities of parking spaces. This module cyclically moves the 3D feature map sequentially in four directions, allowing each pixel to collect global information, such as... Figure 2As shown. Given a three-dimensional feature map tensor X, which contains three dimensions: the number of channels C, the number of rows H, and the number of columns W, global information is gradually aggregated in the vertical and horizontal directions through multiple iterations, and finally each element of the feature map is updated to enhance the expressive power of parking space features. This represents the value of the feature map tensor X at the k-th iteration, where c, i, and j represent the channel, row, and column indices, respectively. The formula for calculating RESA is as follows:

[0069]

[0070] Equation (1) represents the offset step size s in the vertical or horizontal direction at different iteration numbers. k In formulas (2) and (3), L is represented by W and H, respectively. K is the number of iterations. This formula makes the offset step size decrease as the number of iterations increases, so that each pixel of the feature map can aggregate information from a greater distance in each iteration.

[0071]

[0072] Equations (2) and (3) represent the formulas for vertical and horizontal information transmission of feature map X, respectively. This refers to the value of the c-th channel, i-th row, and j-th column of the intermediate feature map Z at the k-th iteration. Z is used to store the aggregated feature information in the vertical or horizontal direction. {m,c,n} It is a set of one-dimensional convolution kernels with a size of N. in ×N out ×w, where N in N out and w represent the number of input channels, the number of output channels, and the core width, respectively. N in and N out Both are equal to C. This indicates that at the k-th iteration, the m-th channel of feature map X, at the (i+s)-th iteration... k modulo H represents the value in the (j+n-1)th column. This formula collects global information for each pixel in the vertical direction.

[0073]

[0074] Equation (4) describes how the elements of the feature map are updated in each iteration, where f is a non-linear activation function. and These represent the original feature map and the updated feature map, respectively. This is represented as an intermediate feature map after linear activation. This formula ensures that the feature map X gradually aggregates global information in each iteration, enhancing the expressive power of the parking space features.

[0075] Specifically, the feature map X is divided into H slices horizontally and W slices vertically. Information is transferred in four directions: "bottom to top" and "top to bottom" for vertical aggregation, and "left to right" and "right to left" for horizontal aggregation. Convolutional kernel weights with the same offset stride are shared across all slices in the same direction. Taking "right to left" as an example, at iteration k=0, s1=1, and the X of each column... i Can receive X i +1 offset feature; at the k=1th iteration, s2=2, X of each column i Can receive X i+2 The offset features. After k iterations, each X i It can aggregate information from the entire feature map, thereby enabling efficient extraction and recognition of parking space features.

[0076] 3) Prediction Network: The output of the RESA module is fed into the prediction network as input for subsequent processing. This prediction network is constructed by connecting four hourglass modules in series. The structure of a single hourglass module is as follows: Figure 3 As shown in -a). The hourglass module, a key component of the network, includes a common sampling module, a downsampling module, an upsampling module, and three output branches, aiming to accurately predict key points on parking lines, parking angle information, and the occupancy status of parking spaces. Each sampling module consists of three convolutional layers. In the first convolutional layer, different types of sampling modules have different parameter configurations based on their functional positioning. Specifically, the common sampling module uses Conv(1 / 0 / 1) convolution operation, the downsampling module uses Conv(3 / 1 / 2) convolution operation, and the upsampling module uses Transposed Conv(3 / 1 / 2) transposed convolution operation. Each module is followed by a PreLU activation function and batch normalization after the convolution operation, as shown in... Figure 3 As shown in -b). These different configurations enable each sampling module to extract specific scales and features from the input data, meeting the network's needs for information processing at different levels. The three output branches are the direct output units of the prediction results. Each output branch consists of an SE attention mechanism, a convolutional layer, and a specific activation function, such as... Figure 3 As shown in -c).

[0077] The main idea of ​​the SE attention mechanism is to improve the model's performance by compressing and activating input features. Specifically, the SE attention mechanism includes two steps: Squeeze and Excitation. In the Squeeze step, the input feature map is compressed into a vector using global average pooling, and then mapped to a smaller vector through a fully connected layer. In the Excitation step, a sigmoid function is used to compress each element of this vector to between 0 and 1, and it is multiplied with the original input feature map to obtain a weighted feature map. The SE attention module is a channel attention module; it can enhance the channel features of the input feature map without changing its size. The formulas for the activation functions Sigmoid and Tanh are as follows.

[0078]

[0079] Where sigmoid(x) is the Sigmoid function with input variable x; tanh(x) is the Tanh function with input variable x.

[0080] The confidence branch uses a combination of SE-Block, Conv, and Sigmoid to output the confidence level of parking space-related information; the angle branch uses a combination of SE-Block, Conv, and Tanh to specifically predict parking space orientation information; and the state branch uses a combination of SE-Block, Conv, and Sigmoid to determine the occupancy status of the parking space. Each branch, through its corresponding structural design, specifically models and predicts different parking space features, ultimately ensuring the comprehensive and accurate output of parking space-related information.

[0081] 4) Knowledge Distillation: This invention introduces the FGD distillation method. This method deeply analyzes the difficulties of knowledge distillation in object detection and innovatively proposes local and global distillation (FGD) methods. Specifically, local distillation separates parking space and background information, guiding the student network to focus on the key pixels and channels of the teacher network, thereby effectively capturing the core information in parking space detection. Global distillation reconstructs the relationships between different pixels and transmits these relationships from the teacher network to the student network to compensate for the global information that may be missing during local distillation. Through the reasonable application of the FGD distillation method, the number of model parameters is reduced from 1.725M to 469.751K while maintaining high accuracy. This ensures that the model can achieve reliable parking space detection function with low resource cost.

[0082] 5.2 Loss Function

[0083] During training, the loss function plays a crucial role. Four loss functions are applied to each output branch of the hourglass module. Details of each loss function are provided below. The output detection head consists of a prediction output head with five features: confidence value (1st branch), angle / direction (2nd branch), and occupancy state (2nd branch). The confidence value determines whether the parking space line keypoint exists; the angle / direction and occupancy state determine the deeper extension direction and occupancy state abstracted from the keypoint, respectively.

[0084] 1) Confidence Loss: Since parking space keypoints occupy a relatively small proportion of pixels in an image, this easily leads to class imbalance, where the number of positive samples (containing parking space keypoints) and negative samples (not containing parking space keypoints) differs significantly. This causes the model to be biased towards the larger proportion of negative samples during training, thus affecting its ability to detect parking space keypoints. To effectively alleviate this problem, this invention uses Focal Loss instead of the standard Cross-Entropy Loss. Focal Loss, by introducing a modulation factor, can adaptively reduce the weight of easily classified samples, making the model pay more attention to difficult-to-classify positive samples. The specific expression is as follows.

[0085] L pos =α pos ∑((y gt -y pred ) 2 log(y pred (7)

[0086]

[0087] Among them, L pos and L neg y represents the confidence loss for positive and negative samples, respectively. pred y is the network's predicted value. gt The label value. The weight coefficient α for positive samples. pos =2, the weighting coefficient α of the negative samples neg =0.1.

[0088] 2) Angle Loss: The two feature output heads of the angle prediction branch are used to determine the directional attributes of key points. Considering the periodicity and continuity of angle information, a cosine-based loss function is used to measure the difference between the predicted angle and the true angle, as shown in the following expression.

[0089]

[0090]

[0091] in, The angles formed by the parking space spacing line with the positive X-axis and positive Y-axis in the pixel coordinate system represent the angles. Based on the constrained angle loss, the angles of the two features can be accurately predicted, and the orientation of the parking space can be inferred in the post-processing algorithm.

[0092] cos pred θ1, cos pred θ2 is the cosine value of the angle predicted by the network, cos gt θ1, cos gt θ2 represents the corresponding label value, and β is the weighting coefficient for the angle loss. θ1 =β θ2 =1.

[0093] 3) Occupancy Status Loss: The occupancy status prediction branch contains two feature output heads, which are used to determine the occupancy status attribute of the parking space. The mean squared error (MSE) loss function is used to measure the difference between the predicted value and the true value, and the specific form is as follows.

[0094] L free =γ free ∑(Free gt -Free pred ) 2 (11)

[0095] L occupied =γ occupied ∑(Occupied gt -Occupied pred ) 2 (12)

[0096] Among them, L free and L occupied These represent the losses for the idle and occupied states, respectively. pred Occupied pred It is a network prediction value, Free gt Occupied gt The corresponding label value. The weighting coefficient γ for the occupied state loss. free =γ occupied =1. The mean squared error loss function can intuitively reflect the degree of deviation between the predicted occupancy status and the actual occupancy status.

[0097] 4) Knowledge distillation loss: Core information in target detection is captured through local distillation, while global distillation compensates for any missing global information during local distillation. The knowledge distillation loss is a weighted sum of the local and global distillation losses, expressed as follows.

[0098] L distillation =δ focal L focal +δglobal L global (13)

[0099] Among them, L distillation For knowledge distillation loss, L focal and L global These represent the local distillation loss and the global distillation loss, respectively, along with the weighting coefficients for the knowledge distillation loss.

[0100] δ focal =δ global =1. In this way, the student network can leverage the advantages of the teacher network, improving its detection performance while maintaining its lightweight structure.

[0101] 5) Total Loss: The total loss function comprehensively considers all the above types of losses and is a weighted sum of them. It also introduces L2 regularization loss to prevent overfitting and enhance the model's generalization ability. Its expression is as follows.

[0102]

[0103] Among them, L total For the total loss, Let be the L2 regularization loss, representing the sum of squares of the parameters in the weight vector. The weight coefficients of the L2 regularization loss are λ = 0.001. The total loss function comprehensively reflects the model's performance at the confidence level (L... pos L neg ),angle Occupied status (L) free L occupied Knowledge distillation (L) distillation The model can continuously adjust its parameters and optimize its performance during training by minimizing the prediction error in terms of L2 regularization and L2 regularization, thereby achieving accurate detection and prediction of parking space-related information.

[0104] 5.3 Parking Space Reasoning

[0105] This invention innovatively proposes a parking space matching algorithm. It synthesizes a real-time surround view image sequence using four fisheye cameras mounted on the vehicle. The transformation matrix from the image coordinate system of the surround view stitched image to the world coordinate system originating from the vehicle's center is pre-calibrated. Once a parking space is detected in the image, its world coordinates can be calculated. The entire parking space detection process is as follows: Figure 4As shown, a pre-trained deep learning model is used to process images, outputting the coordinates, angle, and occupancy status of the parking space's entrance corner. First, the detected entrance line corner and exit line corner are distinguished, and feature matching is performed on the entrance line corner. Second, in the feature point matching stage, feature points are paired based on the actual distance X of the entrance line corner, using distance X to determine the actual distance I of the exit line, and the parking space's exit line corner is calculated (for example, if the distance X between two successfully matched feature points satisfies 100>X>80, the actual distance of the interval line is selected as I1, and the exit line corner is calculated based on the output direction information and I1). Finally, during parking space area fitting, the detected parking space entrance line corner and the calculated exit line corner are combined to form a parking space; furthermore, if the distance between the detected exit line corner and the calculated exit line corner is close, the parking space is directly combined using the detected entrance line corner and these two detected exit line corners. Based on this, the parking space occupancy status is detected.

[0106] 1) Calculation and reasoning of parking space corner points: The parking space angle predicted by the model is used to determine the orientation and extension direction of the parking space, such as... Figure 5 As shown. The model output contains two branches: the first branch is the cosine of the parking space interval line and the X-axis, cosθ1. When cosθ1 equals 0, it is a right angle; when cosθ1 is much greater than 0, it is an acute angle; and when cosθ1 is much less than 0, it is an obtuse angle. This is combined with the usual tilted parking space angles (30 degrees, 45 degrees, etc.) for auxiliary judgment. The second branch is the cosine of the parking space interval line and the Y-axis, cosθ2. When cosθ2 is (0, 1), the parking space interval line extends along the positive Y-axis; when cosθ2 is between (-1, 0), it extends along the negative Y-axis. Finally, the final extension direction is determined based on the specific values. After determining the extension direction, the offset (Δ) of the corner point of the parking space's cutoff line is calculated and inferred using (cosθ1, cosθ2) and the prior distance I. x Δ y And based on the parking space entrance corner point (x, y) and the parking space corner point offset (Δ), x ,Δ y The corner point (x′, y′) of the cutoff line is calculated, and the parking space fitting is completed.

[0107] Δ x =L×cosθ1 (15)

[0108] Δ y =L×cosθ2 (16)

[0109] x′=x+Δ x (17)

[0110] y′=y+Δ y (18)

[0111] Among them, (Δx ,Δ y ) represents the offset of the parking space corner point in the X and Y axes, respectively; is the prior distance, representing the length or width of the parking space; (cosθ1,cosθ2) represents the cosine values ​​of the parking space interval line with the X and Y axes, respectively; (x,y) represents the x and y coordinates of the parking space entrance corner point, respectively; (x′,y′)f represents the x and y coordinates of the cut-off line corner point, respectively.

[0112] 2) Parking space occupancy status: After the parking spaces are fully fitted, the occupancy status is determined based on the entrance corner point of each parking space. Figure 6 For example, the parking space status determination algorithm analyzes the sequence of corner points at the parking space entrance (P0, P1, ..., P...). n-1 The state S i ∈{idle, occupied, mixed}, determine the parking space occupancy status. Each corner point P i State S i The status may be "idle", "occupied", or "mixed". For the mixed status, the idle probability needs to be recorded separately. and occupancy probability Final parking space overall status S total It is determined by the matching results of all adjacent corner point pairs.

[0113] a) Determining the status of adjacent corner points:

[0114]

[0115] Among them, D(P) i P i+1 ) represents the adjacent corner point P i and P i+1 State matching results, S i Corner point P i state, Corner point P i The probability of being idle. Corner point P i The probability of occupancy, and The algorithm first performs state matching on adjacent corner points. If the states of two corner points are the same, that state is directly adopted; if one corner point is in a mixed state, the explicit state of the other corner point is preferred; if both corner points are in a mixed state, the probability is compared to determine the state; otherwise, the state is considered occupied by default.

[0116] b) Dynamic state update rules:

[0117]

[0118] Where D represents the matching status result of adjacent corner points. Corner point P i The probability of being idle. Corner point P i The probability of occupancy. After each match, if the corner point is in a mixed state, its probability value needs to be updated. When the match is "idle", the idle probability is reduced, and if it is exhausted, it switches to the "occupied" state; when the match is "occupied", the occupied probability is reduced, and if it is exhausted, it switches to the "idle" state.

[0119] c) Overall condition of the parking space:

[0120]

[0121] Among them, S total For the overall status of the parking space, D(P) i P i+1 ) represents the adjacent corner point P i and P i+1 The overall status of a parking space is determined by the matching results of all adjacent corner point pairs. Specifically, a parking space is considered vacant only when all matches are "vacant"; otherwise, it is considered occupied. For example, when P0 (vacant) matches P1 (vacant), it is directly determined as a vacant parking space; when P1 (vacant) matches P2 (mixed), the vacant status is prioritized, and the probability value of P2 is updated; subsequently, when P2 changes to an occupied status and matches P3 (mixed), the occupied status is prioritized, and the probability of P3 is updated. This dynamic update mechanism based on probability consumption ensures the accuracy and consistency of status determination.

[0122] 5.4 Experimental Verification

[0123] The method proposed in this invention demonstrates breakthrough performance on two major public datasets, PS2.0 and BODEN: on the PS2.0 dataset, it achieves a precision of 99.46% and a recall of 99.90%; on the BODEN dataset, it also achieves a precision of 95.90% and a recall of 92.70%, which strongly validates the leading advantage of the parking space reasoning model of this invention with quantitative data.

[0124] at the same time, Figure 7 The results visually demonstrate the detection capabilities of this method in extremely complex scenarios. Whether facing challenges arising from complex lighting conditions, diverse parking space types, or blind spots created by vehicles or obstacles obscuring parking spaces, the method of this invention, with its unique algorithm architecture, can accurately locate the outline of parking spaces and efficiently identify their occupancy status. These experimental results fully demonstrate the excellent effectiveness and robustness of this method in complex environments, providing reliable technical support for practical applications.

[0125] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. The various components mentioned in this invention are common technologies in the existing field. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A parking space reasoning model for complex scenarios, characterized in that... Includes the following steps: Step 1: Process the input image using a stacked hourglass network and a RESA (Recurrent Hourglass Shift Aggregator) module to extract contextual features; Step 2: Apply an algorithm based on corner coordinates to match key parking space information, including location, type, and occupancy status; Step 3: Train the lightweight student hourglass module using the local-global distillation (FGD) algorithm to reduce the number of parameters while maintaining generalization ability and detection accuracy.

2. The parking space reasoning model in a complex scenario according to claim 1, characterized in that... The specific steps of step one include: 1) Network adjustment: The input image is compressed by the SPD-Conv layer. Through the joint processing of space-to-depth transformation and non-strut convolution, the output feature map has a resolution of 128×128 and 48 channels. 2) Cyclic Feature Transformation Aggregator: The RESA module utilizes the strong shape prior of parking spaces and the cross-row and column spatial information of pixels to extract features related to parking spaces. This module cyclically moves the features of the 3D feature map in four directions in sequence, so that each pixel collects global information. Given a 3D feature map tensor X, which contains three dimensions: the number of channels C, the number of rows H, and the number of columns W, the module iterates multiple times to gradually aggregate global information in the vertical and horizontal directions, and finally updates each element of the feature map to enhance the expressive power of parking space features.

3. The parking space reasoning model in a complex scenario according to claim 1, characterized in that... The specific steps of step two are as follows: the output of the RESA module is used as the input for subsequent processing and is introduced into the prediction network. The prediction network is constructed by connecting four hourglass modules in series. The hourglass module is a key component of the network and includes a sampling module, a downsampling module, an upsampling module, and three output branches, which are designed to accurately predict key points on the parking line, parking angle information, and the occupancy status of the parking space.

4. The parking space reasoning model in a complex scenario according to claim 3, characterized in that... The same sampling module uses Conv(1 / 0 / 1) convolution operation; the downsampling module uses Conv(3 / 1 / 2) convolution operation; the upsampling module uses Transposed Conv(3 / 1 / 2) transposed convolution operation; and each module is connected to a Prelu activation function and batch normalization after the convolution operation. The three output branches include a confidence branch, an angle branch, and a state branch. Each output branch consists of an SE attention mechanism, a convolutional layer, and a specific activation function.

5. The parking space reasoning model in a complex scenario according to claim 4, characterized in that... The SE attention mechanism consists of two steps: Squeeze and Excitation; The squeezing step is as follows: First, the input feature map is compressed into a vector using a global average pooling operation; then, it is mapped to a smaller vector using a fully connected layer. The excitation step is as follows: Each element of the vector is compressed to between 0 and 1 using a sigmoid function, and then multiplied with the original input feature map to obtain a weighted feature map. The formulas for the activation functions Sigmoid and Tanh are as follows: Where sigmoid(x) is the Sigmoid function with input variable x; tanh(x) is the Tanh function with input variable x.

6. The parking space reasoning model in a complex scenario according to claim 4, characterized in that... The confidence branch uses a combination of SE-Block+Conv+Sigmoid to output the confidence level of parking space-related information; the angle branch uses a combination of SE-Block+Conv+Tanh to specifically predict parking space orientation information; and the state branch uses a combination of SE-Block+Conv+Sigmoid to determine the occupancy status of the parking space.

7. The parking space reasoning model in a complex scenario according to claim 1, characterized in that... The specific steps of step three are as follows: local distillation separates parking space and background information, guides the student network to focus on the key pixels and channels of the teacher network, and captures the core information in parking space detection; global distillation reconstructs the relationship between different pixels and transmits these relationships from the teacher network to the student network to compensate for the global information that may be missing during the local distillation process.

8. The parking space reasoning model in a complex scenario according to claim 7, characterized in that... The joint training of the local-global distillation (FGD) algorithm includes multiple loss functions, specifically confidence loss, angle loss, occupied state loss, knowledge distillation loss, and total loss. The confidence loss uses Focal Loss instead of the standard Cross-Entropy Loss. Focal Loss, by introducing a modulation factor, can adaptively reduce the weight of easily classified samples, making the model pay more attention to difficult-to-classify positive samples. The specific expression is as follows: L pos =α pos ∑((y gt -y pred ) 2 log(y pred )) Among them, L pos and L neg y represents the confidence loss for positive and negative samples, respectively. pred y is the network's predicted value. gt The label value is the weight coefficient α of the positive sample. pos =2, the weighting coefficient α of the negative samples neg =0.1; The angle loss uses a cosine-based loss function to measure the difference between the predicted angle and the true angle, as shown in the following expression: in, The angles formed by the parking space spacing lines with the positive X-axis and positive Y-axis in the pixel coordinate system represent the angles. Based on the constrained angle loss, the angles of these two features can be accurately predicted, and the orientation of the parking space can be inferred in the post-processing algorithm. pred θ1, cos pred θ2 is the cosine value of the angle predicted by the network, cos gt θ1, cos gt θ2 represents the corresponding label value, and β is the weighting coefficient for the angle loss. θ1 =β θ2 =1; The occupancy loss uses the mean squared error (MSE) loss function to measure the difference between the predicted and actual values, and its specific form is as follows: L free =γ free ∑(Free gt -Free pred ) 2 L occupied =γ occupiod ∑(Occupied gt -Occupied pred ) 2 Among them, L free and L occupied These represent the losses for idle and occupied states, respectively. Free pred Occupied pred It is a network prediction, Free gt Occupied gt For the corresponding label value, the weight coefficient γ of the state loss is used. free =γ occupied =1, the mean squared error loss function can intuitively reflect the degree of deviation between the predicted occupancy status and the actual occupancy status; The knowledge distillation loss is a weighted sum of the local distillation loss and the global distillation loss, and its expression is as follows. L distillation =d focal L focal +d global L global Among them, L distillation For knowledge distillation loss, L focal and L global These represent local distillation loss and global distillation loss, respectively, with the weighting coefficient δ for knowledge distillation loss. focal =δ global =1; The total loss function is a weighted sum of confidence loss, angle loss, occupied state loss, and knowledge distillation loss. L2 regularization loss is also introduced to prevent overfitting and enhance the model's generalization ability. Its expression is as follows: Among them, L total For the total loss, Let be the L2 regularization loss, representing the sum of squares of the parameters in the weight vector. The weight coefficients of the L2 regularization loss are λ = 0.

001.

9. The parking space reasoning model in a complex scenario according to claim 1, characterized in that... The specific steps of the corner coordinate algorithm in step two include: 1) Distinguish between the detected entry line corner points and exit line corner points, and perform feature matching on the entry line corner points; 2) In the feature point matching stage, feature points are paired based on the actual distance X of the entrance line corner point. The actual distance I of the cutoff line is determined using distance X, and the corner point of the parking space cutoff line is calculated. 3) When fitting the parking space area, the detected parking space entrance line corner point and the calculated cutoff line corner point are combined to form a parking space; in addition, if the cutoff line corner point detected by the deep learning network model proposed in this invention is close to the calculated cutoff line corner point, the detected parking space entrance line corner point and the two detected cutoff line corner points are directly combined to form a parking space, and on this basis, the parking space occupancy status is detected.