Parking Model Training Method, Parking Method, Computer Device, and Storage Medium
By segmenting and eliminating the training image in obstacle areas and reducing invalid image input, the problems of long training time and many iterations in the existing automatic parking function are solved, and more efficient parking model training is achieved.
Patent Information
- Application Number
- CN202310479698.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-04-28
AI Technical Summary
The existing neural network model with automatic parking function requires long-term training, multiple iterations, and there are a large number of invalid image inputs, resulting in high computing power requirements and excessive time-consuming problems.
By acquiring the training image and segmenting the area according to whether there are obstacles, the obstacle area is eliminated to obtain the training data set, the estimated parking direction is obtained in the initial model, and iteratively trained according to the loss of the driving direction to obtain the parking model.
It effectively reduces the input of invalid image information, reduces the number of iterations of the model, shortens the training time, and improves the computing efficiency.
Smart Images

Figure CN116476853B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of vehicles, and in particular, to a parking model training method, a parking method, a computer device, and a storage medium. Background Art
[0002] With the development of society, in order to meet people's growing material needs, automobiles are often equipped with many intelligent functions, such as an automatic parking function.
[0003] Currently, the automatic parking function can generate a parking path by inputting an image of a parking space area into a neural network model. However, most of the current neural network models are trained based on random image data, and it takes a long time to train and multiple iterations to mature. Moreover, there are a large number of invalid image inputs during use, resulting in problems such as high computing power requirements and excessive time consumption for the automatic parking function. Summary of the Invention
[0004] Based on this, a parking model training method, a parking method, a computer device, and a storage medium are provided to improve the problems of long training time and multiple iterations of the parking model in the prior art.
[0005] On the one hand, a parking model training method is provided, and the method includes:
[0006] Obtain a first training image, and perform region division on the first training image according to the presence or absence of obstacles to obtain a second training image, where the second training image includes an obstacle region and a parking region;
[0007] Eliminate the obstacle region according to the second training image to obtain a training data set, and input the training data set into an initial model to obtain an estimated parking direction;
[0008] Perform iterative training on the initial model according to the loss of the driving direction to obtain a parking model, where the loss of the parking direction is obtained according to the true parking direction and the estimated parking direction, and the true parking direction is obtained according to the second training image.
[0009] In one embodiment, the performing region division on the first training image according to the presence or absence of obstacles to obtain a second training image includes:
[0010] Obtain radar point cloud information of a training environment, where the radar point cloud information includes obstacle point cloud information;
[0011] According to the coincidence matching between the obstacle point cloud information in the radar point cloud information and the first training image, divide the obstacle region and the parking region in the first training image to obtain the second training image.
[0012] In one embodiment, eliminating the obstacle areas based on the second training image to obtain a training dataset includes:
[0013] Dividing the second training image according to a spatial data structure to obtain the training dataset.
[0014] In one embodiment, dividing the second training image according to a spatial data structure to obtain the training dataset includes:
[0015] Obtaining an image matrix based on the second training image, and performing window segmentation on the image matrix to obtain a plurality of blocks;
[0016] Obtaining an indication index of the pixel points within each block, and determining whether the indication index is greater than an obstacle determination threshold;
[0017] Eliminating the blocks with the indication index greater than the obstacle determination threshold, and continuously segmenting the blocks with the indication index less than or equal to the obstacle determination threshold until the number of segmentation times reaches a predetermined number, then determining the remaining blocks as the parking area;
[0018] Obtaining the global coordinates of the pixel points within the parking area as the training dataset.
[0019] In one embodiment, iteratively training the initial model based on the loss of the driving direction includes using a loss function expressed by the following mathematical formula:
[0020]
[0021] where Loss is the loss of the driving direction, y i is the predicted parking direction obtained based on the i-th training dataset, is the true parking direction of the i-th training dataset, and m is the number of the training datasets.
[0022] In one embodiment, iteratively training the initial model based on the loss of the driving direction includes:
[0023] Based on the loss of the driving direction, using the backpropagation algorithm to obtain the error correction amount corresponding to each node in the initial model;
[0024] Updating the connection weights between adjacent two layers of nodes according to the error correction amount of the nodes in the initial model.
[0025] In one embodiment, after inputting the training dataset into the initial model, it further includes:
[0026] Obtaining the time interval between two adjacent inputs, and restarting the initial model when the time interval is greater than a time threshold.
[0027] On the other hand, a parking method is provided, including:
[0028] Obtaining a first environmental image of a parking space environment, performing region differentiation on the first environmental image to obtain a second environmental image, where the second environmental image includes an obstacle region and a parking region, and eliminating the obstacle region according to the second environmental image to obtain an input data set;
[0029] Inputting the input data set into the parking model obtained by the parking model training method to obtain an expected parking direction, so as to park according to the expected parking direction.
[0030] On yet another hand, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the parking model training method is implemented, or the parking method is implemented.
[0031] A computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the parking model training method are implemented, or the steps of the parking method are implemented.
[0032] For the above parking model training method, parking method, computer device, and storage medium, by performing segmentation processing on training samples, the training images are pre-segmented into an obstacle region and a parking region where the vehicle can travel, and then the training data set obtained by eliminating the obstacle region according to the segmented training images is input into the initial model, effectively reducing the input of invalid image information, thereby reducing the number of iterations of the model and shortening the training time. Description of the Drawings
[0033] Figure 1 It is a schematic flowchart of a parking model training method in an embodiment;
[0034] Figure 2 It is a schematic diagram of a parking model in an embodiment;
[0035] Figure 3 It is a schematic flowchart of a parking method in another embodiment;
[0036] Figure 4 It is a structural block diagram of a system architecture in an embodiment;
[0037] Figure 5 It is an internal structure diagram of a computer device in an embodiment. Detailed Embodiments
[0038] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0039] The model in the embodiments of the present application may specifically be a neural network. A neural network may be composed of neural units. By imitating the perception system of the brain neurons, it interprets sensory data in a way of machine perception, marking or clustering the original input data. A neural network may be an operation model, which is composed of a large number of nodes (or called neurons) and the connections between them. A neural unit may refer to an operation unit with Xi and intercept 1 as inputs. The output of this operation unit may be:
[0040]
[0041] where i = 1, 2,..., N, N is a natural number greater than 1, Wi is the weight of Xi, B is the bias of the neural unit. F is the activation function of the neural unit, which is used to perform a non-linear transformation on the features obtained in the neural network, and convert the input signal in the neural unit into an output signal. The output signal of this activation function may be used as the input of the next convolutional layer. The activation function may be the Sigmoid function. A neural network is a network formed by connecting many such single neural units together, that is, the output of one neural unit may be the input of another neural unit. The input of each neural unit may be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field may be a region composed of several neural units.
[0042] The neural network can be applied to the field of autonomous driving. For example, it is used to plan the parking path in a parking environment. At present, path planning is mainly trajectory planning. The trajectory-based planning method mainly uses dynamic programming to solve an optimal global path. However, the disadvantage of this method is that the algorithm takes a long time, the number of iterations of the neural network, that is, the model, is large, and it is particularly dependent on computing power.
[0043] The present application provides a method for training a parking model, which reduces the number of iterations of the neural network model by finding the image regions in a frame of image that may collide with obstacles, thereby improving the operation efficiency.
[0044] In one embodiment, as Figure 1 shown, the method for training the parking model includes the following steps:
[0045] Step 101, obtain a first training image, and perform region segmentation on the first training image according to the presence or absence of obstacles to obtain a second training image, where the second training image includes an obstacle region and a parking region.
[0046] It can be understood that the first training image can be a grayscale image obtained after image preprocessing during the acquisition of parking environment sample data, containing the grayscale value information of the image, or it can be a color image containing multiple image channels. The image recognition system can identify the objects in the first training image through technologies such as image recognition to divide each region in the first training image and obtain a second training image including an obstacle region and a parking region. The second training image contains the position coordinate information of obstacles such as other vehicles, people, walls, and wall columns beside the parking space in the image. All these regions are designated as no-entry regions, while the parking space and the access channel to the parking space in the image serve as the parking region.
[0047] In one embodiment, the first training image is a camera image, and at the same time, the depth value of the training environment is collected through a lidar or a microwave radar, etc., to obtain the radar point cloud information of the training environment. The radar point cloud information includes obstacle point cloud information. According to the coincidence matching between the obstacle point cloud information in the radar point cloud information and the first training image, the corresponding region on the image is matched, and then the size boundary of the obstacle is estimated based on the line of sight, so as to identify the obstacle in the camera image, label the obstacle, distinguish the obstacle region and the parking region in the first training image, and obtain the second training image. To improve the accuracy, the number of matches can be increased.
[0048] The second training image includes the position coordinate information of obstacles such as other vehicles, people, walls, and wall columns beside the parking space in the parking environment.
[0049] Step 102, eliminate the obstacle region according to the second training image to obtain a training data set, and input the training data set into the initial model to obtain an estimated parking direction.
[0050] Exemplarily, the initial model can adopt a Deep Neural Network (DNN), also known as a multi-layer neural network, which can be understood as a neural network with multiple hidden layers. According to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers can be fully connected, that is to say, any neuron in the L-th layer must be connected to any neuron in the L+1-th layer. The DNN neural network has the following linear relationship expression: Y = Σ(W*X + B), where X is the input vector, Y is the output vector, B is the offset vector, W is the weight matrix (weight coefficient), and Σ(·) is the activation function. Each layer simply performs such a simple operation on the input vector X to obtain the output vector Y. Since the DNN has many layers, the number of coefficients W and offset vectors B is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscript corresponds to the index 2 of the output third layer and the index 4 of the input second layer. In summary, the coefficient from the K-th neuron in the L-1-th layer to the J-th neuron in the L-th layer is defined as It should be noted that the input layer does not have the W parameter. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the larger its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network. This embodiment uses a three-layer DNN neural network including an input layer, an intermediate layer, and an output layer as shown in Figure 2 for illustration. The number of nodes in each layer can be determined according to design requirements, and a sufficient number of nodes can be adopted at the beginning. Then, during the training process, according to the relevance of the number of nodes, duplicate nodes can be removed.
[0051] During the implementation process, the second training image can be converted into an image matrix or vector, the obstacle area can be removed, and the training data set Xi~Xn can be obtained and input into the neural network model to obtain the estimated parking direction. It can be understood that the estimated parking direction can be a vector, including the data output of the indication direction and the data output of the magnitude. The output estimated parking direction can also be a trajectory composed of coordinate points.
[0052] Step 103: Iteratively train the initial model based on the loss of the driving direction to obtain a parking model, where the loss of the parking direction is obtained based on the true parking direction and the estimated parking direction, and the true parking direction is obtained based on the second training image.
[0053] During the training process of the DNN network model, each training image can be processed as a training data set as a training sample. A large number of training images are input into the initial model in multiple batches to obtain a model with a high training completion degree. Labels can be added to each sample by means of manual marking or neural network output, etc., to indicate the expected correct output in the sample. In this embodiment, the true parking direction corresponding to each second training image can be noted by means of labeling.
[0054] The loss of the driving direction is obtained by calculating the deviation between the estimated parking direction and the true parking direction. Among them, the greater the loss of the driving direction, the greater the difference between the estimated parking direction output by the initial model and the expected true parking direction. By updating the weight vectors of each layer of neural network in the initial model, the offset compensation ability of the initial model is iteratively trained, so as to gradually reduce the loss of the driving direction and obtain the trained parking model.
[0055] By training the parking model in the above way and performing image segmentation on the training samples, the input of invalid image information is effectively reduced, thereby improving the speed of model training and reducing the number of iterations at the same time.
[0056] As an implementation manner of the above embodiment, obtaining the training data set according to the second training image includes dividing the second training image according to a spatial data structure to obtain the training data set.
[0057] Spatial data structures such as quadtrees, octrees, BVH trees (Bounding Volume Hierarchy BasedOn Tree), etc. are used. Here, the quadtree data structure is taken as an example for illustration. In this embodiment, the lidar and the camera image are overlapped. According to the point cloud information of the obstacles on the lidar, the corresponding area on the image is matched, and then the size boundary of the obstacles is estimated according to the viewing distance. Multiple samples are collected according to this standard, and then the image information is spatially divided according to the quadtree algorithm. The image matrix obtained according to the second training image is divided into multiple different blocks, and the parts with obstacles and the parts without obstacles are separated, and the corresponding quadtree data structure is generated. Traverse the entire quadtree tree structure to check whether there are relevant leaves in the subtree. If so, continue to divide until the tree level reaches the set depth and then stop dividing.
[0058] Exemplarily, the image matrix obtained from the second training image is window-segmented to obtain a plurality of blocks, the indication metrics of the pixel points within each block are obtained, and it is determined whether the indication metrics are greater than the obstacle determination threshold. For example, the indication metrics of the pixel points include the standard deviation of the gray values within each block and the number (or ratio) of pixel points whose gray values are greater than the gray threshold, and the corresponding obstacle determination thresholds include the standard deviation threshold and the number (or ratio) threshold of pixel points. Blocks with a standard deviation greater than the standard deviation threshold and a number (or ratio) of pixel points greater than the number (or ratio) threshold of pixel points are regarded as obstacle areas and removed, and the blocks that do not meet any one or both of the above conditions are further segmented.
[0059] Set the termination conditions for segmentation, such as the number of segmentation times reaching a predetermined number or the gray values of the pixel points within the block all being within a certain threshold range.
[0060] Determine the remaining blocks as the parking area, and obtain the global coordinates corresponding to the pixel points within the parking area (such as the coordinates corresponding in the coordinate system where the three-dimensional object is located) as the training data set. After quadtree segmentation, the input of the training model is very simple position coordinates, and then the input training data are all points that will not collide.
[0061] As Figure 3 shown, the training process of the DNN network model includes the forward propagation and backward propagation processes. In the forward propagation, the training data are sent into the network in batches, and forward calculations are performed layer by layer until the output layer. The output y l of each layer of neural nodes in the forward propagation can be expressed as the following mathematical expression:
[0062] y l = f(a l ) = f(W l Xw -1 + b l )
[0063] where F is the activation function, a l is the output of the linear unit, which is transformed into the output result of the non-linear unit through the activation function. It is determined according to the output X l-1 of the neural nodes in the (L - 1)-th layer and the weight matrix W l of the neural nodes in the L-th layer, and a l is used as the input of the activation function, and b l is the bias of the neural nodes in the L-th layer.
[0064] During the forward propagation process, the output values of all neural nodes can be obtained.
[0065] In this embodiment, a loss function expressed by the following mathematical expression is included:
[0066]
[0067] where Loss is the loss in the driving direction, and y i is the predicted parking direction obtained from the i-th training image, is the true parking direction calibrated for the i-th training image, and m is the training dataset, i.e., the number of training images.
[0068] It can be understood that has the following mathematical expression:
[0069]
[0070] where x i is the input of the input layer, and n is the number of iterations.
[0071] The loss error is calculated using the above loss function to improve the accuracy of the vehicle parking position.
[0072] In the process of backpropagation, according to the chain rule, the backpropagation algorithm is used to obtain the error correction amount corresponding to each node in the initial model, and the gradient of the loss function with respect to each layer is calculated layer by layer. The error correction amount includes the gradient of the loss function with respect to the activation value or the gradient of the loss function with respect to the weight. These are two calculation branches in the backward process. Then, according to the error correction amount of the nodes in the initial model, the connection weights between adjacent layers of nodes are updated.
[0073] Exemplarily, the gradient δ of the loss function of each neural node in the L-th layer with respect to the activation value l can be expressed by the following mathematical expression:
[0074]
[0075] In the above manner, first, δ of the output layer is obtained according to the loss function L , and then δ of each neural node is obtained layer by layer in the reverse direction l .
[0076] The gradient W of the loss function with respect to the weight g can be determined by the following mathematical expression:
[0077]
[0078] Finally, according to the weight gradient obtained in the backward process, the weight is updated:
[0079]
[0080] Among them, η is the learning rate, which is an important hyperparameter in gradient descent. By using the learning rate and the gradient value to update all parameter values, the loss value of the neural network is reduced.
[0081] In one embodiment, it further includes a monitoring process for neural network training. By obtaining the time interval between two adjacent inputs, it is judged whether the time interval is too long. When the time interval is greater than the time threshold, it can be considered that the process has stopped working for a period of time, and the initial model needs to be restarted.
[0082] In one embodiment, a parking method is also provided, as Figure 3 shown, which includes the following steps:
[0083] Step 301: Obtain the first environmental image of the parking space environment, perform region segmentation on the first environmental image to obtain the second environmental image. The second environmental image includes an obstacle region and a parking region, and eliminate the obstacle region according to the second environmental image to obtain the input data set;
[0084] Step 302: Input the input data set into the parking model obtained by the parking model training method described in any one of the above embodiments or implementation manners to obtain the expected parking direction, and park according to the expected parking direction.
[0085] By using the above method, the environmental image of the parking space environment is segmented, reducing the input of invalid image information, effectively reducing the operation time of the pre-collision search degree and vehicle path trajectory correction of the parking model, and outputting the parking direction result faster.
[0086] As an implementation manner of the parking method, the obtaining the second environmental image by performing region segmentation on the first environmental image includes:
[0087] Obtain the radar point cloud information of the parking space environment, and the radar point cloud information includes obstacle point cloud information; according to the coincidence matching of the obstacle point cloud information in the radar point cloud information and the first environmental image, segment the obstacle region and the parking region in the first environmental image to obtain the second environmental image.
[0088] It should be noted that the process of obtaining the second environmental image can be completed in one branch of the parking model, or in another model, and then input into the parking model of the present application. The present application does not make any restrictions on this.
[0089] In one embodiment, the method for obtaining the input data set according to the second environmental image includes dividing the second environmental image according to the spatial data structure to obtain the input data set.
[0090] Exemplarily, it includes:
[0091] An environmental image matrix is obtained according to the second environmental image, and the environmental image matrix is window-divided to obtain a plurality of blocks;
[0092] An indication index of pixel points in each block is obtained, and it is determined whether the indication index is greater than an obstacle determination threshold;
[0093] The blocks with the indication index greater than the obstacle determination threshold are eliminated, and the blocks with the indication index less than or equal to the obstacle determination threshold are continuously divided until the number of divisions reaches a predetermined number, and then the remaining blocks are determined as the parking area;
[0094] The global coordinates of pixel points in the parking area are obtained as the input data set.
[0095] It should be understood that although Figure 1 、 Figure 3 each step in the flowchart of Figure 1 、 Figure 3 is displayed in sequence according to the indication of the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0096] See Figure 4 , an embodiment of the present application provides a system architecture. As shown in the system architecture, the data acquisition device 460 can be used to acquire training data. After the data acquisition device 460 acquires the training data, these training data are stored in the database 430, and the training device 420 trains to obtain the target model / rule 401 based on the training data maintained in the database 430.
[0097] The following describes how the training device 420 obtains the target model / rule 401 based on the training data. Exemplarily, the training device 420 processes multiple frames of sample images to output corresponding prediction labels, calculates the loss between the prediction labels and the original labels of the samples, and updates the classification network based on this loss until the prediction labels are close to the original labels of the samples or the difference between the prediction labels and the original labels is less than a threshold, thereby completing the training of the target model / rule 401. For specific descriptions, refer to the training method in this article.
[0098] The target model / rule 401 in the embodiments of this application can specifically be a neural network. It should be noted that in actual applications, the training data maintained in the database 430 does not necessarily all come from the collection of the data collection device 460, and it may also be received from other devices. Additionally, it should be noted that the training device 420 does not necessarily train the target model / rule 401 entirely based on the training data maintained in the database 430, and it may also obtain training data from the cloud or other places for model training. The above descriptions should not be regarded as limitations on the embodiments of this application.
[0099] The target model / rule 401 trained according to the training device 420 can be applied to different systems or devices, such as being applied to Figure 4 the execution device 420 shown. The execution device 420 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, augmented reality (AR) / virtual reality (VR), a vehicle-mounted terminal, a television, etc., or it can also be a server or the cloud, etc. In Figure 4 it, the execution device 420 is configured with a transceiver 412, and this transceiver can include an input / output (I / O) interface or other wireless or wired communication interfaces, etc., for data interaction with external devices. Taking the I / O interface as an example, a user can input data to the I / O interface through the client device 440.
[0100] During the preprocessing of the input data by the execution device 420, or during the relevant processing such as the execution of calculations by the computing module 412 of the execution device 420, the execution device 420 can call the data, code, etc. in the data storage system 450 for corresponding processing, or can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 450.
[0101] Finally, the I / O interface 412 returns the processing result to the client device 440, so as to provide it to the user.
[0102] It is worth noting that the training device 420 can generate corresponding target models / rules 401 based on different training data for different targets or different tasks. The corresponding target models / rules 401 can be used to achieve the above targets or complete the above tasks, so as to provide the required results for the user.
[0103] In the appendix Figure 4In the case shown, the user can manually input data, and this manual input can be operated through the interface provided by the transceiver 412. In another case, the client device 440 can automatically send input data to the transceiver 412. If the client device 440 is required to automatically send input data and user authorization is needed, the user can set the corresponding permissions in the client device 440. The user can view the results output by the execution device 420 on the client device 440, and the specific presentation form can be display, sound, movement, etc. The client device 440 can also be used as a data acquisition end to collect the input data input to the transceiver 412 and the output results of the transceiver 412 as shown in the figure as new sample data and store them in the database 430. Of course, it is also possible not to collect data through the client device 440, but directly store the input data input to the transceiver 412 and the output results of the transceiver 412 as shown in the figure as new sample data in the database 430.
[0104] It should be noted that Figure 4 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in Figure 4 , the data storage system 450 is an external memory relative to the execution device 420. In other cases, the data storage system 450 can also be placed in the execution device 420.
[0105] As Figure 4 shown, a target model / rule 401 is trained according to the training device 420. The target model / rule 401 can be a parking model in the embodiments of the present application.
[0106] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it is used to implement a parking model training method or a parking method. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0107] Those skilled in the art can understand that Figure 5 the structure shown in Figure 5 is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0108] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0109] Obtain a first training image, and perform region segmentation on the first training image according to the presence or absence of obstacles to obtain a second training image, where the second training image includes an obstacle region and a parking region;
[0110] Eliminate the obstacle region according to the second training image to obtain a training data set, and input the training data set into an initial model to obtain an estimated parking direction;
[0111] Iteratively train the initial model according to the loss of the driving direction to obtain a parking model, where the loss of the parking direction is obtained according to the true parking direction and the estimated parking direction, and the true parking direction is obtained according to the second training image.
[0112] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0113] Obtain the radar point cloud information of the training environment, where the radar point cloud information includes obstacle point cloud information;
[0114] According to the coincidence matching of the obstacle point cloud information in the radar point cloud information and the first training image, segment the obstacle region and the parking region in the first training image to obtain the second training image.
[0115] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0116] Divide the second training image according to the spatial data structure to obtain the training data set.
[0117] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0118] The iterative training of the initial model according to the loss of the driving direction includes using a loss function expressed by the following mathematical formula:
[0119]
[0120] where Loss is the loss in the driving direction, and y i is the predicted parking direction obtained according to the i-th training dataset, is the true parking direction of the i-th training dataset, and m is the number of the training datasets.
[0121] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0122] According to the loss in the driving direction, use the backpropagation algorithm to obtain the error correction amount corresponding to each node in the initial model;
[0123] According to the error correction amount of the nodes in the initial model, update the connection weights between adjacent two layers of nodes.
[0124] In one embodiment, when the processor executes the computer program, the following steps are further implemented:
[0125] Obtain a first environmental image of the parking space environment, perform region segmentation on the first environmental image to obtain a second environmental image, the second environmental image includes an obstacle region and a parking region, and eliminate the obstacle region according to the second environmental image to obtain an input dataset;
[0126] Input the input dataset into the parking model obtained according to the parking model training method described in any one of the foregoing embodiments to obtain an expected parking direction, so as to park according to the expected parking direction.
[0127] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0128] Obtain a first training image, perform region segmentation on the first training image according to the presence or absence of obstacles to obtain a second training image, the second training image includes an obstacle region and a parking region;
[0129] Eliminate the obstacle region according to the second training image to obtain a training dataset, input the training dataset into the initial model to obtain a predicted parking direction;
[0130] Iteratively train the initial model according to the loss in the driving direction to obtain a parking model, where the loss in the parking direction is obtained according to the true parking direction and the predicted parking direction, and the true parking direction is obtained according to the second training image.
[0131] In one embodiment, when the computer program is executed by the processor, the following steps are further implemented:
[0132] Obtain the radar point cloud information of the training environment, where the radar point cloud information includes obstacle point cloud information;
[0133] According to the coincidence matching between the obstacle point cloud information in the radar point cloud information and the first training image, segment the obstacle area and the parking area in the first training image to obtain the second training image.
[0134] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0135] Divide the second training image according to the spatial data structure to obtain the training dataset.
[0136] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented:
[0137] The iterative training of the initial model according to the loss of the driving direction includes using a loss function expressed by the following mathematical formula:
[0138]
[0139] where Loss is the loss of the driving direction, y i is the estimated parking direction obtained according to the i-th training dataset, is the true parking direction of the i-th training dataset, and m is the number of the training datasets.
[0140] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0141] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0142] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A parking model training method, characterized in that, Including: Obtain a first training image, and perform region differentiation on the first training image according to the presence or absence of obstacles to obtain a second training image, where the second training image includes an obstacle region and a parking region; Eliminate the obstacle region according to the second training image to obtain a training data set, and input the training data set into an initial model to obtain an estimated parking direction, including dividing the second training image according to a spatial data structure to obtain the training data set; Iteratively train the initial model according to the loss of the driving direction to obtain a parking model, where the loss of the parking direction is obtained according to the true parking direction and the estimated parking direction, and the true parking direction is obtained according to the second training image; The step of dividing the second training image according to a spatial data structure to obtain the training data set includes: Obtain an image matrix according to the second training image, and perform window segmentation on the image matrix to obtain a plurality of blocks; Obtain an indication index of the pixel points in each block, and determine whether the indication index is greater than an obstacle judgment threshold; Eliminate the blocks with the indication index greater than the obstacle judgment threshold, and continue to segment the blocks with the indication index less than or equal to the obstacle judgment threshold until the number of segmentation times reaches a predetermined number, then determine the remaining blocks as the parking region; Obtain the global coordinates of the pixel points in the parking region as the training data set.
2. The parking model training method according to claim 1, characterized in that, The iterative training of the initial model according to the loss of the driving direction includes using a loss function expressed by the following mathematical formula: where Loss is the loss in the driving direction, and y i is the predicted parking direction obtained from the i-th training dataset, is the true parking direction of the i-th training dataset, and m is the number of the training datasets.
3. The parking model training method according to claim 1, wherein The iterative training of the initial model according to the loss of the driving direction includes: According to the loss of the driving direction, use the backpropagation algorithm to obtain an error correction amount corresponding to each node in the initial model; Update the connection weights between adjacent layers of nodes according to the error correction amount of the nodes in the initial model.
4. The parking model training method according to claim 1, characterized in that After inputting the training data set into the initial model, it further includes: Obtain the time interval between two adjacent inputs, and when the time interval is greater than a time threshold, restart the initial model.
5. A parking method, characterized in that, Including: Obtain a first environmental image of the parking space environment, perform region differentiation on the first environmental image to obtain a second environmental image, where the second environmental image includes an obstacle region and a parking region, and eliminate the obstacle region according to the second environmental image to obtain an input data set, including dividing the second environmental image according to a spatial data structure to obtain the input data set, where the step of dividing the second environmental image according to a spatial data structure to obtain the input data set includes: Obtain an environmental image matrix according to the second environmental image, and perform window segmentation on the environmental image matrix to obtain a plurality of blocks; Obtain an indication index of the pixel points in each block, and determine whether the indication index is greater than an obstacle judgment threshold; Eliminate the blocks with the indication index greater than the obstacle judgment threshold, and continue to segment the blocks with the indication index less than or equal to the obstacle judgment threshold until the number of segmentation times reaches a predetermined number, then determine the remaining blocks as the parking region of the second environmental image; Obtain the global coordinates of the pixel points in the parking region of the second environmental image as the input data set; Input the input data set into the parking model obtained by the parking model training method described in any one of claims 1-4 to obtain the expected parking direction, so as to park according to the expected parking direction.
6. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the parking model training method described in any one of claims 1 to 4, or implements the parking method described in claim 5.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the parking model training method described in any one of claims 1 to 4, or implements the steps of the parking method described in claim 5.
Citation Information
Patent Citations
Method for reducing neural network training sample size based on image segmentation
CN110689057A
Trajectory tracking control method, device and system for automatic parking and storage medium
CN114834441A
Automatic parking method and system based on neural network
CN115303263A