Parking space detection model training method, parking space detection method and device
By introducing image detection auxiliary heads and image segmentation auxiliary heads into the parking space detection model, more parking space related information is obtained, and the problem of low performance in the existing technology of parking space detection models relying on local information is solved, achieving a more efficient parking space detection effect.
Patent Information
- Application Number
- CN202311586969.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The existing parking space detection model only relies on local parking space information during the training process, resulting in low parking space detection performance and poor detection effect.
By acquiring the parking space area image, using the parking space detection model for feature extraction, combining the image detection auxiliary head and the image segmentation auxiliary head, more parking space-related information is obtained from the feature maps output from different downsampling layers, a more effective target loss function is determined, and then iterative training is carried out.
The parking space detection performance of the parking space detection model has been improved, the accuracy of the parking space detection effect has been significantly improved, and the parking space-related information can be more comprehensively characterized.
Smart Images

Figure CN120047916A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent parking, and specifically relates to a training method for a parking space detection model, a parking space detection method, and a device therefor. Background Art
[0002] With the development of intelligent driving technology, Automated Valet Parking (AVP) has also developed rapidly. Parking space detection is the core content of the AVP project. The accuracy of parking space detection directly affects the path planning and parking decision-making for later assisted parking. Therefore, effective parking space detection is crucial. In related technologies, a trained parking space detection model is used to perform the parking space detection task to achieve the purpose of parking space detection. However, during the training process of the parking space detection model for performing the parking space detection task, only local information of the parking space (such as the corner points of the parking space entrance, etc.) is relied on for model training, and a lot of useful parking space-related information is not effectively utilized, resulting in low parking space detection performance of the trained parking space detection model and low accuracy of the parking space detection effect. Summary of the Invention
[0003] In view of this, the present invention provides a training method for a parking space detection model, a parking space detection method, and a device therefor to solve the problems of low parking space detection performance and poor parking space detection effect existing in related technologies.
[0004] In a first aspect, the present invention provides a training method for a parking space detection model, and the training method includes:
[0005] Obtain an image of a parking space area, where the image of the parking space area is an image in a training set;
[0006] Use the parking space detection model to perform feature extraction on the image of the parking space area to obtain a first feature map, and use an image detection auxiliary head to obtain a second feature map from the result output by the first downsampling layer of the parking space detection model, and use an image segmentation auxiliary head to obtain a third feature map from the result output by the second downsampling layer of the parking space detection model; wherein, both the image detection auxiliary head and the image segmentation auxiliary head are feature detection heads added to the parking space detection model;
[0007] Determine the result of a target loss function according to the first feature map, the second feature map, and the third feature map;
[0008] Based on the result of the target loss function, perform iterative training on the parking space detection model.
[0009] Based on the first feature map output by the parking space detection model, the present invention also obtains a second feature map through an image detection auxiliary head and a third feature map through an image segmentation auxiliary head. The first feature map, the second feature map, and the third feature map can comprehensively represent more parking space-related information. Therefore, in the process of training the parking space detection model, the present invention uses more parking space-related information, not limited to using local parking space information, but also using overall parking space information. The present invention can determine a more effective target loss function result based on the above image processing results, thereby improving the parking space detection performance of the trained parking space detection model and helping to significantly improve the accuracy of the parking space detection effect.
[0010] In an alternative embodiment, determining the result of the target loss function according to the first feature map, the second feature map, and the third feature map includes:
[0011] Determining the result of the first loss function according to the first feature map and multiple second feature maps, and determining the result of the second loss function according to multiple third feature maps;
[0012] Superposing the result of the first loss function and the result of the second loss function to obtain the target loss function.
[0013] The present invention determines the result of the first loss function according to the first feature map and the second feature map and determines the result of the second loss function according to the third feature map, realizing the combination of the object detection task and the semantic segmentation task, paying attention to more information in the process of training the parking space detection model, and improving the performance of the parking space detection model.
[0014] In an alternative embodiment, determining the result of the second loss function according to multiple third feature maps includes:
[0015] Fusing multiple third feature maps to obtain a fourth feature map; the third feature map includes parking space boundary line features;
[0016] Determining the result of the second loss function according to the fourth feature map.
[0017] The present invention can also fuse the feature maps of different layers after adding the label information of the segmentation task into a new layer based on the image segmentation auxiliary head and add it to the network, helping the parking space detection model to focus on the required information, and learning more information such as clearly visible parking space boundary lines through the image segmentation auxiliary head, improving the accuracy of parking space recognition.
[0018] In an alternative embodiment, superposing the result of the first loss function and the result of the second loss function to obtain the target loss function includes:
[0019] Assigning a first weight to the result of the first loss function and a second weight to the result of the second loss function;
[0020] Determine the first product of the result of the first loss function and the first weight, and determine the second product of the result of the second loss function and the second weight, and use the sum of the first product and the second product as the target loss function.
[0021] The present invention can determine the target loss function based on the first loss function, the first weight, the second loss function and the second weight, and further realize paying attention to more information during the training process of the parking space detection model, so as to improve the performance of the parking space detection model.
[0022] In an alternative embodiment, determining the result of the target loss function according to the first feature map, the second feature map and the third feature map includes:
[0023] Calculate the target loss function based on the differences between the first feature map, the second feature map and the third feature map and the corresponding label information respectively;
[0024] Wherein, the first feature map is used to represent the overall information of the predicted parking space, the second feature map is used to represent the first local information of the predicted parking space, and the third feature map is used to represent the second local information of the predicted parking space.
[0025] When calculating the target loss function of the present invention, an accurate result of the target loss function is calculated based on the first feature map representing global information, the second feature map and the third feature map representing local information.
[0026] In an alternative embodiment, the multiple first downsampling layers and the multiple second downsampling layers are the same.
[0027] The present invention sets the multiple first downsampling layers and the multiple second downsampling layers in the same way, which can reduce the difficulty of the training process of the vehicle detection model.
[0028] In an alternative embodiment, the parking space detection model is a deep residual network.
[0029] The present invention can effectively train the deep residual network, so as to obtain a parking space detection model with better parking space recognition effect.
[0030] In a second aspect, the present invention provides a parking space detection method, and the detection method includes:
[0031] Collect a target image of a target area, where the target area includes an area of a parking space to be detected;
[0032] Input the target image into the trained parking space detection model obtained by the training method of the parking space detection model according to any embodiment of the present invention, so as to output a plurality of detection frames through the parking space detection model;
[0033] After post-processing by non-maximum suppression, the target detection box is selected from multiple detection boxes; the target detection box is used to represent the detection result of the parking space.
[0034] The parking space detection model of the present invention can pay attention to the overall information and local information of the parking space, so as to obtain more comprehensive parking space-related features, and thus can predict more appropriate multiple detection boxes, select more accurate target detection boxes, and improve the accuracy of the parking space detection effect.
[0035] In a third aspect, the present invention provides a training device for a parking space detection model, and the training device includes:
[0036] An image acquisition module, configured to acquire an image of a parking space area, and the image of the parking space area is an image in the training set;
[0037] A feature extraction module, configured to extract features from the image of the parking space area by using the parking space detection model to obtain a first feature map. The feature extraction module is further configured to obtain a second feature map from the result output by the first downsampling layer of the parking space detection model by using an image detection auxiliary head, and the feature extraction module is further configured to obtain a third feature map from the result output by the second downsampling layer of the parking space detection model by using an image segmentation auxiliary head; wherein, both the image detection auxiliary head and the image segmentation auxiliary head are feature detection heads added to the parking space detection model.
[0038] A loss determination module, configured to determine the result of the target loss function according to the first feature map, the second feature map, and the third feature map;
[0039] An iterative training module, configured to perform iterative training on the parking space detection model based on the result of the target loss function.
[0040] In a fourth aspect, the present invention provides a parking space detection device, and the parking space detection device includes:
[0041] An acquisition module, configured to acquire a target image of a target area, and the target area includes an area of a parking space to be detected;
[0042] A detection module, configured to input the target image into the trained parking space detection model obtained by the training method of the parking space detection model according to any embodiment of the present invention, so as to output multiple detection boxes through the parking space detection model;
[0043] A screening module, configured to screen out the target detection box from multiple detection boxes by post-processing of non-maximum suppression; the target detection box is used to represent the detection result of the parking space. Description of the Drawings
[0044] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a schematic flowchart of a method for training a parking space detection model according to an embodiment of the present invention;
[0046] Figure 2 It is a schematic flowchart of a method for training another parking space detection model according to an embodiment of the present invention;
[0047] Figure 3 It is a schematic diagram of obtaining a second feature map from the result output by the first downsampling layer of the parking space detection model by using an image detection auxiliary head according to an embodiment of the present invention;
[0048] Figure 4 It is a schematic diagram of obtaining a fourth feature map after fusing multiple third feature maps according to an embodiment of the present invention;
[0049] Figure 5 It is a schematic diagram of obtaining the target loss function after superimposing the results of the first loss function and the results of the second loss function according to an embodiment of the present invention;
[0050] Figure 6 It is a schematic flowchart of a parking space detection method according to an embodiment of the present invention;
[0051] Figure 7 It is a structural block diagram of a training device for a parking space detection model according to an embodiment of the present invention;
[0052] Figure 8 It is a structural block diagram of a parking space detection device according to an embodiment of the present invention;
[0053] Figure 9 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Specific Embodiments
[0054] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0055] In the related art, the methods for parking space detection are generally divided into three categories, namely object detection, semantic segmentation, and Transformer queries. Among them, object detection is further divided into parking space corner detection, parking space anchor box detection, and parking space line segment detection; considering the actual resources and complex parking scenarios, the parking space corner detection method is used the most. For the parking space corner detection algorithm based on object detection, the target key points are set as the two entrance corners of the parking space. The reason for choosing the entrance corner as the target key point for detection is that the entrance corner is usually relatively clear and is located at the intersection of the entrance line and the parking space boundary line. However, it is difficult to accurately obtain the angle of the parking space boundary line by regressing based on two corners in the existing parking space detection model, and the deviation of the predicted parking space line angle is relatively large, resulting in a low detection accuracy of the current detection method based on the parking space entrance corner, which seriously affects the subsequent SLAM (Simultaneous Localization and Mapping) tracking and parking planning tasks. Therefore, it is crucial to improve the accuracy of the parking space detection results of the parking space detection model.
[0056] According to an embodiment of the present invention, there is provided an embodiment of a training method for a parking space detection model. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0057] In this embodiment, a training method for a parking space detection model is provided, which can be used in a computer device. Figure 1 It is a flowchart of the training method for the parking space detection model according to an embodiment of the present invention, as Figure 1 shown, and this process includes the following steps:
[0058] Step S101, obtain an image of the parking space area, and the image of the parking space area is an image in the training set.
[0059] Among them, the training set can include multiple batches of images of the parking space area, and each batch of images of the parking space area is respectively used to train the parking space detection model.
[0060] Step S102, use the parking space detection model to extract features from the image of the parking space area to obtain a first feature map, and use the image detection auxiliary head to obtain a second feature map from the result output by the first downsampling layer of the parking space detection model, and use the image segmentation auxiliary head to obtain a third feature map from the result output by the second downsampling layer of the parking space detection model; wherein, both the image detection auxiliary head and the image segmentation auxiliary head are feature detection heads added to the parking space detection model.
[0061] Specifically, this embodiment can use the backbone in the parking space detection model to extract features from the parking space area image. The first feature map, the second feature map, and the third feature map in this embodiment can respectively include not only the parking space corner point features, the parking space entrance line features, and the parking space category features, but also, without limitation, the parking space boundary line features. The parking space boundary line features can specifically refer to judging the category to which the parking space belongs by calculating the angle between the boundary line and the entrance line.
[0062] A feature map refers to the output image or feature matrix obtained after passing through the backbone network. The backbone network is usually a deep learning network composed of multiple convolutional layers and pooling layers. The stacking of these convolutional layers and pooling layers can gradually transform the original image into a higher-level feature representation; the earlier layers will extract low-level features such as edges and textures, while the deeper layers will extract more high-level semantic features such as the shape, parts, and relationships of the object.
[0063] Among them, the parking space detection model includes a backbone network and a multi-task feature detection head connected in sequence. The multi-task feature detection head refers to the feature detection heads for object detection tasks and semantic segmentation tasks. The image detection auxiliary head in this embodiment is used for the feature extraction process of the object detection task, and the image segmentation auxiliary head is used for the feature extraction process of the semantic segmentation task. It should be understood that the image detection auxiliary head, the image segmentation auxiliary head, and the multi-task feature detection head are all feature detection heads. The feature detection head (Feature Detection Head) is a network structure component in deep learning for feature extraction and the generation of feature maps. The feature detection head includes, for example, a global average pooling layer, a fully connected layer, a convolutional layer, etc., for transforming the high-dimensional backbone network features into more advanced and finer-grained feature representations. The global average pooling layer is used to reduce the spatial dimension of the feature map and retain the features in the channel dimension; the fully connected layer is used to compress and abstract the features to generate task-related feature representations; the convolutional layer can be used to further extract local features and patterns.
[0064] In some alternative embodiments, the parking space detection model is a deep learning network, specifically a deep residual network.
[0065] The deep residual network in this embodiment can be Resnet18 (Deep Residual Network 18). The deep learning network used in the present invention is not limited to Resnet18. Figure 3 and Figure 4The parking space detection model involved herein is illustrated by taking Resnet18 as an example. In this embodiment, through feature map visualization, it is found that the earlier layers of the deep residual network are more likely to focus on overall information, such as straight lines and color blocks, while the later layers are more focused on detailed information, such as key points; the information focused on by different convolutional layers often varies. For example, the parking space frame border line can be clearly observed in conv1 (the first convolutional layer), the parking space frame border line becomes less obvious in conv4 (the fourth convolutional layer), and the parking space frame border line becomes very discontinuous in conv8 (the eighth convolutional layer), and it more tends to find the features of the cross intersection. The present invention can extract more convolutional layers containing the required information according to the task requirements to guide the model training. The embodiment of the present invention can effectively train the deep residual network, so as to obtain a parking space detection model with better parking space recognition effect.
[0066] In some alternative embodiments, the multiple first downsampling layers and the multiple second downsampling layers are the same.
[0067] Combined Figure 3 and Figure 4 As shown, taking Resnet18 as an example, both the multiple first downsampling layers and the multiple second downsampling layers can be three different downsampling layers. Figure 3 Three image detection auxiliary heads (aux_head) are shown. The outputs calculated from the input image data in these three downsampling layers are respectively extracted and added to the final loss calculation, which helps the model converge faster and more accurately through backpropagation. Since the receptive field sizes of the three downsampling layers are different, this embodiment can obtain three different types of feature information, significantly improving the knowledge richness of the model. Setting the multiple first downsampling layers and the multiple second downsampling layers in the same way can reduce the difficulty in the training process of the vehicle detection model.
[0068] In this embodiment, feature maps are extracted in the downsampling stage, the receptive field size will change, and the network will focus on information at different levels. By adding the feature map information extracted from different downsampling layers to the total loss calculation (such as the loss of angle regression), the deep learning network in the training is adaptively adjusted.
[0069] Step S103, determine the result of the target loss function according to the first feature map, the second feature map, and the third feature map.
[0070] Specifically, the image detection auxiliary head, the image segmentation auxiliary head, and the multi-task feature detection head can respectively convert the underlying feature maps of the corresponding networks into prediction results corresponding to the tasks, and then compare these prediction results with the corresponding ground truth labels to calculate the corresponding target loss functions.
[0071] In some alternative embodiments, the above step S103 includes:
[0072] Calculate the target loss function based on the differences between the first feature map, the second feature map, and the third feature map and their corresponding label information respectively; wherein, the first feature map is used to represent the overall information of the predicted parking space, the second feature map is used to represent the first local information of the predicted parking space, and the third feature map is used to represent the second local information of the predicted parking space.
[0073] Among them, the overall information of the parking space is the global information of the parking space; the global information of the parking space includes, for example, the parking space category, the occupancy situation of the parking space, and the approximate location, etc. The first local information and the second local information are both local information of the parking space, and the local information of the parking space includes, for example, the angle of the parking space boundary line, the position of the parking space entrance corner point, etc.
[0074] Combined with the three image detection auxiliary heads shown in the foregoing embodiments Figure 3 Calculate the results of the loss functions corresponding to each image detection auxiliary head respectively, and corresponding weights can be set for the results of each loss function, and adjust them according to the order of magnitude of the loss curves of each detection head drawn, so that the order of magnitude of the results of the loss functions of all detection heads remains at a similar level to avoid the loss function from oscillating and causing the non-convergence of the model; similar to the process of adjusting the weights of single-task loss functions by multi-task loss functions, in this embodiment, all task weights can be preset to the same value, and then adjust the order of magnitude of all loss function curves through multiple experiments for visualization.
[0075] For the image segmentation auxiliary head, the label information corresponding to the third feature map is specifically a mask (MASK) map obtained based on the manual annotation method. Since the label information has a clear direction set manually, it can guide the backbone network to actively learn accurate parking space boundary line feature information.
[0076] In this embodiment, when calculating the target loss function of the parking space detection model, a more accurate result of the target loss function is calculated based on the first feature map representing the global information and the second and third feature maps representing the local information.
[0077] Step S104, based on the result of the target loss function, perform iterative training on the parking space detection model.
[0078] Specifically, the result of the target loss function can be used to measure the difference between the prediction result and the label, and update the parameters of the parking space detection model through the backpropagation algorithm. In each iterative training, the parking space detection model will perform forward propagation calculation according to the given training data, then calculate the error according to the calculation result and the true label of the training data, and then perform backpropagation to update the model parameters, so that the parking space detection model gradually converges and optimizes.
[0079] In this embodiment, the first feature map output by the parking space detection model, the second feature map obtained through the image detection auxiliary head, and the third feature map obtained through the image segmentation auxiliary head are used. The first feature map, the second feature map, and the third feature map can comprehensively represent more parking space-related information. Therefore, in the process of training the parking space detection model in this embodiment, more parking space-related information is used, not limited to using local parking space information, but also using the overall parking space information. This embodiment can determine a more effective target loss function result based on the above image processing results, and further improve the parking space detection performance of the trained parking space detection model, which helps to significantly improve the accuracy of the parking space detection effect. In this embodiment, by adding an image detection auxiliary head and an image segmentation auxiliary head during the training stage of the deep learning network, the result of the target loss function gradually decreases during the training process, that is, it can help the deep learning network converge faster, and the performance of the model gradually improves, obtaining a parking space detection model with better performance. In addition, this embodiment has less dependence on prior knowledge during the model training process and can be applied to more scenarios.
[0080] In this embodiment, a training method for a parking space detection model is provided, which can be used in a computer device. Figure 2 It is a flowchart of the training method for the parking space detection model according to an embodiment of the present invention, as Figure 2 shown, and this process includes the following steps:
[0081] Step S201, obtain an image of the parking space area, and the image of the parking space area is an image in the training set. For details, please refer to Figure 1 step S101 of the embodiment shown here, which will not be elaborated here.
[0082] Step S202, use the parking space detection model to extract features from the image of the parking space area to obtain a first feature map, and use the image detection auxiliary head to obtain a second feature map from the result output by the first downsampling layer of the parking space detection model, and use the image segmentation auxiliary head to obtain a third feature map from the result output by the second downsampling layer of the parking space detection model; wherein, both the image detection auxiliary head and the image segmentation auxiliary head are feature detection heads added to the parking space detection model. For details, please refer to Figure 1 step S102 of the embodiment shown here, which will not be elaborated here.
[0083] Step S203, determine the result of the target loss function according to the first feature map, the second feature map, and the third feature map.
[0084] Specifically, the above step S203 includes:
[0085] Step S2031, determine the result of the first loss function according to the first feature map and multiple second feature maps, and determine the result of the second loss function according to multiple third feature maps.
[0086] Taking the calculation of the angle between the side line and the entrance line as an example, the result of the first loss function is denoted as loss1 (Loss 1), and this embodiment calculates it in the following manner:
[0087]
[0088] where j xpred represents the component value in the x direction predicted based on the first feature map and multiple second feature maps, and j xgt represents the true component value in the x direction, and j ypred represents the component value in the y direction predicted based on the first feature map and multiple second feature maps, and j ygt represents the true component value in the y direction.
[0089] In this embodiment, the multi-task feature detection head, the image detection auxiliary head, and the image segmentation auxiliary head are all angle detection heads.
[0090] Combined with Figure 3 As shown, the embodiment of the present invention adds feature maps of different layers to the total loss calculation based on the image detection auxiliary head to help the parking space detection model learn more knowledge. Compared with the related art where only the information obtained from the last layer of the network can be extracted, resulting in the information of the intermediate layers of the network being ignored and lost, this embodiment can combine different features of more layers through the image detection auxiliary head, and the feature map can be obtained by passing the input data through the required layers in sequence by a specific method of the used framework (such as pytorch) to obtain the output feature map.
[0091] In some alternative embodiments, determining the result of the second loss function according to the multiple third feature maps includes:
[0092] Step a1, fusing the multiple third feature maps to obtain a fourth feature map; the third feature maps include parking space side line features.
[0093] Among them, the parking space side line features of this embodiment include the parking space side line angle, and the parking space side line angle specifically refers to the angle between the parking space side line and the entrance line.
[0094] Step a2, determining the result of the second loss function according to the fourth feature map.
[0095] Specifically, the result of the second loss function is determined based on the difference between the fourth feature map and the corresponding label information. Taking Resnet18 as an example, the MASK map of the segmentation task is added to the input data as the label information (ground truth) for network training; in this embodiment, the feature information of three downsampling layers is extracted from the Resnet backbone network with the new MASK map added; the three downsampling feature maps of different sizes are fused into one output layer through convolution operations and added to the loss calculation as an auxiliary detection head.
[0096] Combined Figure 4 As shown, based on the image segmentation auxiliary head, the feature maps of different layers after adding the label information of the segmentation task are fused into a new layer and added to the network to help the parking space detection model focus on the required information. The clear visible parking space boundary line can be learned through the image segmentation auxiliary head to assist in detecting the true angle of the parking space line.
[0097] In this embodiment, two different detection heads, namely the image detection auxiliary head and the image segmentation auxiliary head, are used as aids to help the deep learning network learn the information related to the parking space boundary line more accurately. Specifically, two auxiliary detection heads added in the multi-task feature detection head part can help improve the angle detection accuracy. Thus, on the basis of combining semantic segmentation and object detection, the detection accuracy of the parking space boundary line angle is improved with less prior knowledge, that is, the parking space detection performance of the parking space detection model is improved and the parking space detection effect is enhanced.
[0098] Step S2032: The results of the first loss function and the second loss function are superimposed to obtain the target loss function.
[0099] Among them, the way of superimposing the result of the first loss function and the result of the second loss function in this embodiment can be weighted addition or direct addition, etc.
[0100] As Figure 5 shown, the parking space detection annotation file represents an image annotated with the position and / or bounding box of the parking space. A set of densely distributed candidate boxes or anchor boxes can be generated on the image through the feature detection head grid matrix and then input into the backbone network. The segmentation task annotation file can represent the annotation file in the image segmentation task, which is used to provide the label information of the segmentation task. The mask map is the visual representation of the segmentation task annotation file. In this embodiment, the mask map is specifically input into the backbone network. Figure 5 The feature detection head in Figure 3 may include a multi-task feature detection head (the feature detection head shown in Figure 5 and an image detection auxiliary head, which is used to calculate loss1. The image segmentation auxiliary head in is used to calculate loss2, and then the overall loss is calculated through loss1 and loss2.
[0101] In some alternative embodiments, step S2032 includes:
[0102] Step b1: Assign a first weight to the result of the first loss function and a second weight to the result of the second loss function.
[0103] Specifically, for the result loss1 (loss 1) of the first loss function, assign a first weight weight1 (weight 1); for the result loss2 (loss 2) of the second loss function, assign a second weight weight2 (weight 2).
[0104] Step b2: Determine the first product of the result of the first loss function and the first weight, and determine the second product of the result of the second loss function and the second weight, and use the sum of the first product and the second product as the target loss function.
[0105] Specifically, the first product is weight1 * loss1, and the second product is weight2 * loss2.
[0106] As Figure 5 shown, and in combination with Figure 3 , the result of the second loss function can be expressed as loss2 (loss 2), and the final loss = weight1 * loss1 + weight2 * loss2. Among them, the specific values of weight 1 (weight1) and weight 2 (weight2) are set according to the actual situation. During the parameter adjustment stage, first try in a 1:1 ratio, and then fine-tune in two directions respectively. For example, each time it can be adjusted up and down by 10% until the multiple loss curves of the parking space detection model drawn tend to be stable. After adding two types of auxiliary detection heads, the present invention largely solves the problem of large deviation in the prediction of the angle of the parking space boundary line caused by relying on two corner points for regression. At the same time, due to the addition of the annotation file for the segmentation task, it helps the backbone network learn specified features and can also help improve the overall detection accuracy of the parking space. In this embodiment, determining the result of the first loss function according to the first feature map and the second feature map and determining the result of the second loss function according to the third feature map realizes the combination of the object detection task and the semantic segmentation task, so that more information can be concerned during the training process of the parking space detection model, and the performance of the parking space detection model can be improved.
[0107] Step S204: Based on the result of the target loss function, perform iterative training on the parking space detection model. For details, please refer to Figure 1 step S104 of the embodiment shown, which will not be elaborated here.
[0108] In this embodiment, a parking space detection method is provided, which can be used in a computer deviceFigure 6 is a flowchart of a parking space detection method according to an embodiment of the present invention. As Figure 6 shown, the process includes the following steps:
[0109] Step S601, collect a target image of a target area, where the target area includes an area of a parking space to be detected.
[0110] For the image of the area of the parking space to be detected, it can be a spliced image, for example, an image spliced from multiple images collected by a multi-channel camera (such as a four-channel camera).
[0111] Step S602, input the target image into a trained parking space detection model obtained by a training method based on a parking space detection model, so as to output multiple detection frames through the parking space detection model.
[0112] Among them, the parking space detection model is specifically a parking space detection model obtained by the training method in any embodiment of the present invention. Specifically, a feature map is extracted through a backbone network in the parking space detection model, and then a multi-task feature detection head generates multiple detection frames according to the feature map extracted by the backbone network.
[0113] Among them, the trained parking space detection model includes a backbone network and a multi-task feature detection head connected in sequence.
[0114] Step S603, through a non-maximum suppression post-processing method, screen out a target detection frame from multiple detection frames; the target detection frame is used to represent the detection result of the parking space.
[0115] The non-maximum suppression (NMS) post-processing method in this embodiment removes redundant candidate frames (multiple detection frames), so as to obtain a more accurate detection frame. For example, the following method can be adopted: select a confidence threshold: according to the confidence of the predicted detection frame, set a threshold, and this threshold will be used to screen the detection frame; sort in descending order of confidence: according to the confidence of the detection frame, sort the predicted detection frames in descending order of confidence; select the detection frame with the highest confidence: select the detection frame with the highest confidence from the sorted detection frame list and add it to the final output list, that is, obtain the detection result of the parking space.
[0116] Based on the parking space detection model trained with the image detection auxiliary head and the image segmentation auxiliary head, the parking space detection model in this embodiment can focus on the overall information and local information of the parking space, identify more useful information in the target image, thereby obtaining more comprehensive parking space-related features. Based on the more comprehensive parking space-related features, it is possible to predict more appropriate multiple detection frames, screen out more accurate target detection frames, and improve the accuracy of the parking space detection effect. During the intelligent parking process, based on the parking space detection method provided in this embodiment, the possibility of successful parking can be greatly improved, the user experience can be enhanced, and the user satisfaction is higher.
[0117] In this embodiment, a training device for a parking space detection model is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0118] This embodiment provides a training device for a parking space detection model, as Figure 7 shown, including:
[0119] An image acquisition module 701, configured to acquire an image of a parking space area, and the image of the parking space area is an image in the training set.
[0120] A feature extraction module 702, configured to extract features from the image of the parking space area by using the parking space detection model to obtain a first feature map. The feature extraction module 702 is further configured to obtain a second feature map from the result output by the first downsampling layer of the parking space detection model by using the image detection auxiliary head. The feature extraction module 702 is further configured to obtain a third feature map from the result output by the second downsampling layer of the parking space detection model by using the image segmentation auxiliary head; wherein, both the image detection auxiliary head and the image segmentation auxiliary head are feature detection heads added to the parking space detection model.
[0121] A loss determination module 703, configured to determine the result of the target loss function according to the first feature map, the second feature map, and the third feature map.
[0122] An iterative training module, configured to perform iterative training on the parking space detection model based on the result of the target loss function.
[0123] In some optional implementation manners, the loss determination module 703 includes:
[0124] A determination unit, configured to determine the result of the first loss function according to the first feature map and multiple second feature maps, and determine the result of the second loss function according to multiple third feature maps.
[0125] An overlay unit for overlaying the results of the first loss function and the results of the second loss function to obtain a target loss function.
[0126] In some alternative embodiments, the determination unit includes:
[0127] A fusion subunit for fusing a plurality of third feature maps to obtain a fourth feature map; the third feature maps include parking space boundary line features.
[0128] A determination subunit for determining the result of the second loss function according to the fourth feature map.
[0129] In some alternative embodiments, the overlay unit includes:
[0130] An assignment subunit for assigning a first weight to the result of the first loss function and for assigning a second weight to the result of the second loss function.
[0131] A calculation subunit for determining a first product of the result of the first loss function and the first weight and for determining a second product of the result of the second loss function and the second weight, and for taking the sum of the first product and the second product as the target loss function.
[0132] In some alternative embodiments, the loss determination module 703 is specifically configured to calculate a target loss function based on the differences between the first feature map, the second feature map, and the third feature map and the corresponding label information; wherein, the first feature map is used to represent the overall information of the predicted parking space, the second feature map is used to represent the first local information of the predicted parking space, and the third feature map is used to represent the second local information of the predicted parking space.
[0133] In some alternative embodiments, the plurality of first downsampling layers and the plurality of second downsampling layers are the same.
[0134] In some alternative embodiments, the parking space detection model is a deep residual network.
[0135] In this embodiment, a parking space detection device is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0136] This embodiment provides a parking space detection device, as Figure 8 shown, including:
[0137] An acquisition module 801 for acquiring a target image of a target area, the target area including an area of a parking space to be detected.
[0138] The detection module 802 is configured to input a target image into a trained parking space detection model obtained by the training method of the parking space detection model according to any embodiment of the present invention, so as to output a plurality of detection frames through the parking space detection model.
[0139] The screening module 803 is configured to screen out a target detection frame from the plurality of detection frames by means of non-maximum suppression post-processing; the target detection frame is used to represent the detection result of the parking space.
[0140] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding embodiments above, and will not be repeated here.
[0141] The training device and / or the parking space detection device of the parking space detection model in this embodiment are presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0142] The embodiment of the present invention further provides a computer device having the above-mentioned Figure 7 shown training device of the parking space detection model and / or Figure 8 shown parking space detection device.
[0143] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 9 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 9 In
[0144] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field-programmable gate array, a generic array logic, or any combination thereof.
[0145] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0146] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may further include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0147] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memories.
[0148] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 40 may be connected through a bus or other means. Figure 9 Taking connection through a bus as an example.
[0149] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (such as an LED), and a tactile feedback device (such as a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.
[0150] Embodiments of the present invention also provide a computer-readable storage medium. The method according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0151] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A training method for a parking space detection model, It is characterized in that The method comprises: Acquire a parking area image, where the parking area image is an image in a training set; Using a parking space detection model to perform feature extraction on the parking space area image to obtain a first feature map, and using an image detection auxiliary head to obtain a second feature map from a result output by a first downsampling layer of the parking space detection model, and using an image segmentation auxiliary head to obtain a third feature map from a result output by a second downsampling layer of the parking space detection model; wherein the image detection auxiliary head and the image segmentation auxiliary head are both feature detection heads added to the parking space detection model; Determine a result of a target loss function according to the first feature map, the second feature map, and the third feature map; Based on the result of the target loss function, the parking space detection model is iteratively trained.
2. The method according to claim 1, It is characterized in that The determining a result of a target loss function according to the first feature map, the second feature map, and the third feature map includes: Determine a result of a first loss function according to the first feature map and the plurality of the second feature maps, and determine a result of a second loss function according to the plurality of the third feature maps; The result of the first loss function and the result of the second loss function are superimposed to obtain the target loss function.
3. The method according to claim 2, It is characterized in that The step of determining a result of a second loss function according to the plurality of third feature maps comprises: The plurality of third feature maps are merged to obtain a fourth feature map; the third feature map includes a parking space edge feature; A result of the second loss function is determined according to the fourth feature map.
4. The method according to claim 2, It is characterized in that The superimposing the result of the first loss function and the result of the second loss function to obtain the target loss function includes: Assigning a first weight to a result of the first loss function, and assigning a second weight to a result of the second loss function; Determine a first product of a result of the first loss function and the first weight, and determine a second product of a result of the second loss function and the second weight, and use the sum of the first product and the second product as the target loss function.
5. The method according to claim 1, It is characterized in that The determining a result of a target loss function according to the first feature map, the second feature map, and the third feature map includes: Calculating a target loss function based on differences between the first feature map, the second feature map, and the third feature map and corresponding label information; The first feature map is used to characterize the overall information of the predicted parking space, the second feature map is used to characterize the first local information of the predicted parking space, and the third feature map is used to characterize the second local information of the predicted parking space.
6. The method according to any one of claims 1 to 5, It is characterized in that The plurality of first downsampling layers and the plurality of second downsampling layers are the same.
7. The method according to any one of claims 1 to 5, It is characterized in that The parking space detection model is a deep residual network.
8. A parking space detection method, It is characterized in that The method comprises: Acquire a target image of a target area, wherein the target area includes an area of a parking space to be detected; Inputting the target image into a parking space detection model that has been trained based on the training method according to any one of claims 1 to 7, so as to output a plurality of detection frames through the parking space detection model; A target detection frame is screened out from the multiple detection frames by a non-maximum suppression post-processing method; the target detection frame is used to represent the detection result of the parking space.
9. A training device for a parking space detection model, It is characterized in that The device comprises: An image acquisition module is used to acquire a parking area image, where the parking area image is an image in a training set; A feature extraction module, used to extract features from the parking space area image using a parking space detection model to obtain a first feature map. The feature extraction module is also used to obtain a second feature map from a result output by a first downsampling layer of the parking space detection model using an image detection auxiliary head. The feature extraction module is also used to obtain a third feature map from a result output by a second downsampling layer of the parking space detection model using an image segmentation auxiliary head. The image detection auxiliary head and the image segmentation auxiliary head are both feature detection heads added to the parking space detection model. A loss determination module, configured to determine a result of a target loss function according to the first feature map, the second feature map, and the third feature map; An iterative training module is used to iteratively train the parking space detection model based on the result of the target loss function.
10. A parking space detection device, It is characterized in that The device comprises: An acquisition module, used for acquiring a target image of a target area, wherein the target area includes an area of a parking space to be detected; A detection module, configured to input the target image into a parking space detection model that has been trained based on the training method according to any one of claims 1 to 7, so as to output a plurality of detection frames through the parking space detection model; The screening module is used to screen out a target detection frame from the multiple detection frames by non-maximum suppression post-processing; the target detection frame is used to represent the detection result of the parking space.