Perception model training method, system and equipment and storage medium
By constructing a multi-task bird's-eye view perception model and optimizing training with image and point cloud data, the problems of data truth accuracy and multi-task collaboration in the automatic parking perception model were solved, achieving efficient environmental perception and rapid iteration, and meeting automotive-grade deployment requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-14
AI Technical Summary
Existing automatic parking perception models suffer from problems such as insufficient accuracy of training data, poor coordination of multi-task perception performance, and high data supplementation costs, resulting in low accuracy in automatic parking environment perception and an inability to quickly iterate and optimize.
A multi-task bird's-eye view perception model is constructed. By combining two-dimensional image data acquired by the image acquisition module and point cloud data acquired by the point cloud acquisition module, high-quality training data is generated through inverse perspective transformation and semi-automatic annotation. The perception model is then optimized to achieve multi-task collaborative training.
It achieves efficient collaborative training of multi-task perception for automatic parking, improves the model's environmental perception accuracy and rapid iteration capability, reduces data supplementation costs, and meets automotive-grade deployment requirements.
Smart Images

Figure CN121861606A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of automatic parking technology, specifically to training methods, systems, devices, and storage media for perception models. Background Technology
[0002] Automated parking systems are a core function of advanced driver assistance systems (ADAS), and their performance hinges on the accuracy of the environmental perception module. Perception models are evolving from handling single tasks to integrating multiple tasks to provide a more comprehensive understanding of the environment. However, most existing perception models and training methods are designed for single tasks or simply pieced together multi-task approaches. Therefore, a training scheme for perception models capable of handling multiple tasks is needed. Summary of the Invention
[0003] This invention provides a training method, system, device, and storage medium for a perception model to achieve efficient collaborative training of multi-task perception for automatic parking.
[0004] Firstly, a training method for a perception model is provided, applied to automatic parking of a vehicle, the vehicle including an image acquisition module; the training method includes: Acquire target image data and target perception model; wherein, the target image data is two-dimensional image data of the vehicle in the target parking scene acquired by the image acquisition module; the target perception model is a pre-built multi-task bird's-eye view perception model; The target perception model is trained and optimized based on image data to obtain a trained target perception model; the trained target perception model is used to perceive multiple target tasks for automatic parking.
[0005] In some embodiments, the target perception model includes a backbone network, a feature pyramid network, a first detection head, a second detection head, and a third detection head; training and optimizing the target perception model based on the target image data to obtain a trained target perception model includes: The target image data is subjected to feature extraction based on the backbone network to obtain a first feature map; wherein, the first feature map includes semantic features and positional features; The semantic features and the location features are fused using a feature pyramid network to obtain a second feature map. The first detection head, the second detection head, and the third detection head are trained and optimized based on the second feature map to obtain a trained target perception model. The multiple target tasks include a first target task, a second target task, and a third target task. The first detection head is used to output the first target task, the second detection head is used to output the second target task, and the third detection head is used to output the third target task.
[0006] In some embodiments, the first objective task is to detect and classify static targets; the first detection head includes a convolutional network, a first output head, and a second output head; the method for training and optimizing the first detection head based on a second feature map includes: The second feature map is input into the convolutional network, and the target class probability map and bounding box regression map are output through the first output head and the second output head, respectively. The parameters of the target category probability map and bounding box regression map are optimized according to the first preset loss function so that the static target can be detected and classified by the optimized first detection head.
[0007] In some embodiments, the second objective task is to detect a dynamic target and predict the motion state of the dynamic target; the second detection head includes a first encoder; the method for training and optimizing the second detection head based on a second feature map includes: The second feature map of a preset number of consecutive frames is input into the first encoder; The first encoder captures the correlation between dynamic targets in adjacent frames to track and estimate the speed of the dynamic targets. The dynamic target is predicted to belong to the interval and its offset relative to the midpoint of the interval based on the preset binning regression algorithm. Determine the heading angle of a dynamic target based on the interval and offset; The velocity and heading angle are optimized according to the second preset loss function so that the dynamic target can be detected by the optimized second detection head and the motion state of the dynamic target can be predicted.
[0008] In some embodiments, the third objective task is to identify a safe driving area; the third detection head includes a second encoder and a first decoder; the method for training and optimizing the third detection head based on the second feature map includes: The target segmentation map is output based on the second encoder, the first decoder, and the second feature map. Based on the target segmentation map and the preset line-surface joint learning algorithm, the binary segmentation map of the safe driving area and the edge map of the obstacle are output. The parameters of the binary segmentation map and edge map are optimized according to the third preset loss function so that the safe driving area can be identified by the optimized third detection head.
[0009] In some embodiments, after acquiring the target image data, the method further includes: The target image data is labeled and projected using a preset labeling tool and a preset inverse perspective transformation algorithm to obtain the bird's-eye view planar coordinate representation of the target image data in the vehicle coordinate system; And generate a target format annotation file based on the target image data represented by the planar coordinates of the bird's-eye view.
[0010] In some embodiments, the vehicle further includes a point cloud acquisition module; the training method further includes: In the early stages of training the target perception model, point cloud data is acquired by the point cloud acquisition module, and the target perception model is trained and optimized based on the point cloud data.
[0011] Secondly, a training system for a perception model is also provided, applied to the automatic parking of a vehicle, the vehicle including an image acquisition module; the training system includes: The first acquisition module is used to acquire target image data; wherein, the target image data is two-dimensional image data of the vehicle in the target parking scene acquired by the image acquisition module; The second acquisition module is used to acquire the target perception model; the target perception model is based on a multi-task bird's-eye view perception model. The training module is used to train and optimize the target perception model based on the target image data to obtain the trained target perception model; the trained target perception model is used to perceive multiple target tasks.
[0012] Thirdly, an electronic device is also provided, including a memory and a processor, wherein a computer program is stored in the memory, and the computer program, when executed by the processor, implements the method described in the first aspect.
[0013] Fourthly, a computer-readable storage medium is also provided, on which a computer program is stored, the computer program being loaded by a processor to perform the steps of the method described in the first aspect.
[0014] Beneficial Effects: This application provides a training method, system, device, and storage medium for a perception model. The training method includes: acquiring target image data and a target perception model; wherein the target image data is two-dimensional image data of a vehicle in a target parking scene acquired by an image acquisition module; the target perception model is a pre-constructed multi-task bird's-eye view-based perception model; training and optimizing the target perception model based on the target image data to obtain a trained target perception model; wherein the trained target perception model is used to perceive multiple target tasks of automatic parking. The perception model training method provided in this application achieves efficient collaborative training of multi-task perception for automatic parking by constructing a multi-task bird's-eye view-based perception model and using two-dimensional image data of a vehicle in a target parking scene acquired by an image acquisition module to train and optimize the multi-task perception model. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of a solution based on 2D image annotation and BEV perception model; Figure 2 This is a schematic diagram of a perception scheme based on pure LiDAR point clouds; Figure 3 This is a flowchart of a training method for a perception model provided in an embodiment of this application; Figure 4 This is a schematic diagram of the overall process of training the perception model provided in the embodiments of this application; Figure 5 This is a schematic diagram of the point cloud map construction and annotation process provided in the embodiments of this application; Figure 6 This is a schematic diagram of the 2D data supplementation process provided in the embodiments of this application; Figure 7 This is a schematic diagram of the training process of the multi-task BEV perception model provided in the embodiments of this application; Figure 8 This is an overall flowchart of the multi-task BEV perception model training method provided in the embodiments of this application; Figure 9 This is a schematic diagram of the principle structure of a training system for a perception model provided in the embodiments of this application. Detailed Implementation
[0017] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0018] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified.
[0019] "A and / or B" includes the following three combinations: A only, B only, and a combination of A and B.
[0020] The use of "applies to" or "configured to" in this application implies open and inclusive language, which does not exclude the applicability to or configuration to devices performing additional tasks or steps. Additionally, the use of "based on" implies openness and inclusivity, because processes, steps, calculations, or other actions "based on" one or more of the stated conditions or values may in practice be based on additional conditions or values beyond those stated.
[0021] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0022] The applicant's research revealed that automated parking systems are a core function of advanced driver assistance systems (ADAS), and their performance hinges on the accuracy of the environmental perception module. Currently, vision-based bird's-eye view (BEV) perception technology has become the mainstream approach for parking perception due to its ability to provide intuitive bird's-eye view scene representations. The implementation of this technology relies on two main elements: high-quality training data ground truth and efficient neural network models. Training data typically originates from precise measurements of the real world, initially relying primarily on 2D image annotations, but this has inherent limitations due to the lack of spatial information. LiDAR can provide accurate 3D point clouds, but it is costly. Simultaneously, perception models are evolving from handling single tasks to integrating multiple tasks to provide a more comprehensive environmental understanding. However, how to acquire high-precision ground truth data at low cost and how to design a model architecture that can efficiently coordinate multiple tasks remain critical technical challenges that need to be overcome in this field.
[0023] There are two main architectures for related multi-task perception models. The first approach is based on 2D image annotation and BEV perception models. Figure 1 This is a schematic diagram of a solution based on 2D image annotation and a BEV perception model. For example... Figure 1 As shown, this scheme first acquires 2D images of the parking scene using an onboard camera, and then annotators directly mark the bounding boxes or outlines of targets such as parking spaces and obstacles on the images. Subsequently, using camera calibration parameters, the images and annotation information are transformed into the BEV space through inverse perspective mapping to form training samples. The model typically employs a Convolutional Neural Network (CNN) or Transformer architecture, with one or more simple detection heads connected to the end, and uses common cross-entropy loss and smooth L1 loss for end-to-end training. When uncovered scenes occur, a ground truth vehicle is dispatched to perform a complete set of data acquisition and annotation. Its components and functions are: 1. 2D data acquisition module (camera): acquires scene images; 2. 2D manual annotation module: generates initial ground truth; 3. Inverse Perspective Mapping (IPM) transformation module: maps image coordinates to the BEV plane; 4. BEV perception model: performs target detection and segmentation.
[0024] The second approach is a perception scheme based on pure LiDAR point clouds. Figure 2 This is a schematic diagram of a perception scheme based on pure LiDAR point clouds. (Example) Figure 2As shown, this scheme relies on 3D point cloud data acquired by LiDAR. Targets (e.g., parking space point cloud clusters, obstacles) are directly labeled on the point cloud, or the point cloud is voxelized to generate BEV feature maps. The model primarily focuses on the detection and segmentation of static targets, with a relatively simple network structure and a lack of dedicated design for temporal modeling of dynamic targets. This scheme is a closed system and does not incorporate low-cost visual data as a supplementary channel. Its components and functions are: 1. Point cloud acquisition module (LiDAR): acquiring 3D environmental information; 2. Point cloud labeling module: generating accurate ground truth values in 3D space; 3. Point cloud processing and model: performing perceptual reasoning based on 3D data.
[0025] However, both of these approaches have certain drawbacks. The first approach has the following disadvantages: 1. Low accuracy of training data. This is because the proposed method uses annotation techniques on 2D images with perspective effects and occlusion. Parking scenarios, however, are characterized by low viewing angles and multiple occlusions, inevitably making it difficult for annotators to accurately determine the true geometric boundaries of the targets. The causal chain is: Cause (technical method) - 2D annotation affected by perspective and occlusion - Effect (technical defect): Inherent bias in annotation - biased data in model learning, reducing the upper limit of perceptual accuracy.
[0026] 2. The model suffers from poor coordination and low efficiency in multi-task perception performance. This is because the proposed solution uses a generic multi-task head and loss function that is not differentiated for the characteristics of different tasks such as static targets, dynamic targets, and drivable areas in parking scenarios. However, different perception tasks have fundamentally different requirements for feature representation and optimization objectives. The causal chain is: cause (technical means) - "one-size-fits-all" network structure and loss function - effect (technical defects): tasks interfere with each other, making it difficult to achieve optimal performance simultaneously, resulting in low model computational efficiency.
[0027] 3. High cost and slow response to mass production data replenishment. This is because the solution establishes a closed data replenishment process that "relies on a truth collection vehicle when encountering new scenarios," and the deployment of the truth collection vehicle is costly and time-consuming. The causal chain is: cause (technical means) - rigid and expensive data replenishment link - effect (technical defects): slow model iteration, unable to quickly respond to mass production needs.
[0028] The disadvantage of the second option is: 1. Weak dynamic target perception capability and incomplete functionality. This is because: since this scheme mainly focuses on static environment modeling, its model architecture lacks a dedicated design for time-series tracking of dynamic targets (such as time-series fusion networks or trajectory prediction heads). The causal chain is: cause (technical means) - limited model function design - effect (technical defects): unable to effectively perceive and predict dynamic targets such as pedestrians and vehicles, resulting in insufficient system security.
[0029] 2. Limited data source, limited generalization ability, and high cost. This is because: as this solution is a closed system relying solely on LiDAR, it does not provide a low-cost data supplementation mechanism. When LiDAR has detection blind spots or needs to supplement data for specific scenarios, there is no effective alternative. The causal chain is: cause (technical means) - limited data source and rigid cost - effect (technical defects): limited coverage of training dataset, insufficient model generalization ability, and high overall data cost.
[0030] In summary, there are three key bottlenecks in the development and mass production application of automatic parking perception models, which restrict further improvement and rapid iteration of model performance. These are as follows: First, the accuracy of the ground truth in the training data is insufficient: the mainstream method relies on manual annotation of 2D images. In typical parking scenarios (such as narrow garages), due to severe occlusion by obstacles, perspective distortion, and blurred imaging of distant targets, the geometric position of the labeled bounding boxes or contours has inherent biases. Training the model with such biased data directly reduces the reliability of parking environment perception and poses safety hazards.
[0031] Secondly, the models suffer from limited perception capabilities and low efficiency: existing perception models mostly employ single-task or simply pieced-together multi-task heads, making it difficult to coordinate and optimize the three core capabilities required for automated parking—high-precision static target (parking space, pillar) detection, continuous tracking of dynamic targets (pedestrians, vehicles), and pixel-level drivable area segmentation. The lack of targeted design for the network structure and loss function of each task results in a difficulty in balancing accuracy and real-time performance, failing to meet automotive-grade deployment requirements.
[0032] Third, the cost and cycle of supplementing mass production data are high: When the model needs to be adapted to new scenarios (such as irregular parking spaces or temporary obstacles) during the mass production stage, the existing solution relies heavily on "truth value acquisition vehicles" equipped with expensive sensors such as LiDAR to re-collect data and label the entire process. This method has high hardware costs, cumbersome operation procedures, and data output cycles that can last for several weeks, resulting in the model being unable to respond quickly to market feedback and passively slowing down iterative optimization.
[0033] In view of this, embodiments of this application provide a training method, system, device, and storage medium for a perception model. Embodiments of this application construct a multi-task bird's-eye view-based perception model and use two-dimensional image data of the vehicle in the target parking scene acquired by the image acquisition module to train and optimize the multi-task perception model, thereby achieving efficient collaborative training of multi-task perception for automatic parking.
[0034] Figure 3This is a flowchart illustrating a training method for a perception model provided in this embodiment. On one hand, this embodiment provides a training method for a perception model, applicable to any stage of the perception model in a vehicle control system (including the development stage, testing stage, and practical application stage), achieving efficient collaborative training of multi-task perception for automatic parking. This method can be executed by a perception model training system, which can be implemented in software and / or hardware, and can be configured in the processor or controller of the vehicle control system. Please refer to... Figure 3 The method includes the following steps: Step 110: Obtain target image data and target perception model.
[0035] The target image data consists of two-dimensional image data of the vehicle in the target parking scenario acquired by the image acquisition module; the target perception model is a pre-built perception model based on multi-task bird's-eye view (BEV).
[0036] The image data acquisition module is a common vehicle camera, such as a mass-produced vehicle surround-view fisheye camera or a front-view camera.
[0037] In some embodiments, the vehicle further includes a point cloud acquisition module; the training method of the perception model further includes: in the early stage of training the target perception model, acquiring point cloud data collected by the point cloud acquisition module, and training and optimizing the target perception model based on the point cloud data.
[0038] The point cloud acquisition module collects raw point cloud data from high-precision sensors. Examples of high-precision sensors include lidar, GPS, and inertial measurement units (IMUs).
[0039] Figure 4This is a schematic diagram of the overall process of the training method for the perception model provided in this application embodiment. Exemplarily, the training method for this perception model mainly includes three core parts: point cloud map construction and annotation, 2D image data supplementation, and multi-task BEV perception model and training. These parts work together to form a complete closed loop from data preparation and model training to practical application. Point cloud map construction and annotation refers to constructing the three-dimensional ground truth of the real scene using point cloud data collected by high-precision sensors in the early stages of multi-task BEV perception model development. 2D image data supplementation refers to rapid data supplementation and annotation of 2D image data collected by ordinary vehicle-mounted cameras for specific problems or new scenarios during the mass production and optimization stages of the multi-task BEV perception model. The multi-task BEV perception model and training is the core algorithm of this application, used to receive visual input and simultaneously complete multiple perception tasks in the BEV space. See also... Figure 4 The training workflow of this perception model is as follows: First, a basic ground truth dataset is constructed using point clouds for the initial training of the multi-task BEV perception model. Then, during the testing phase (i.e., the optimization phase) and the mass production phase, new scene data is continuously collected and labeled using 2D image data for iterative optimization of the multi-task BEV perception model. Finally, the trained and optimized multi-task BEV perception model is generated and deployed to the vehicle to achieve real-time parking environment perception.
[0040] The mass production phase of the multi-task BEV perception model refers to the stage where the model has completed multiple rounds of optimization before mass production, meets mass production standards, is successfully installed on mass-produced vehicles and delivered to users, and enters the stage of real-vehicle operation, continuous data collection, and feedback. The optimization phase of the multi-task BEV perception model includes not only post-mass production optimization but also "iterative optimization before mass production" and "continuous optimization after mass production." This includes iterative optimization after the initial model training is completed and before mass production, addressing issues exposed during testing (such as missing scenarios in laboratory and road tests). It also includes continuous optimization after mass production based on market feedback and testing issues, ultimately supplementing the training with new data to create a new model, which is then synchronized to the vehicle via Over-the-Air (OTA) technology.
[0041] Figure 5 This is a schematic diagram illustrating the point cloud map construction and annotation process provided in this application embodiment. For example, see [link to relevant documentation]. Figure 5 In the early stages of training the target perception model, the specific implementation principle and process of acquiring point cloud data collected by the point cloud acquisition module and training and optimizing the target perception model based on the point cloud data are as follows: mainly including two stages: point cloud map construction and point cloud annotation.
[0042] Point cloud map construction involves processing raw point cloud data collected by sensors such as LiDAR, GPS / IMU, and others mounted on a ground truth acquisition vehicle. The specific process of point cloud map construction includes the following steps: Step 1: Preprocess the raw point cloud data. This preprocessing includes point cloud denoising (such as removing outliers) and point cloud filtering.
[0043] Step 2: Combine the pose information provided by high-precision GNSS / IMU to perform registration of point clouds in consecutive frames.
[0044] High-precision GNSS / IMU typically integrates GPS with Real-Time Kinematic (RTK) positioning technology and an IMU. The specific process of registering point clouds across consecutive frames using the pose information provided by the high-precision GNSS / IMU is as follows: The pose information provided by the high-precision GNSS / IMU is the core initial constraint for point cloud registration. Essentially, it maps the relative coordinate point clouds acquired by the lidar to a unified global coordinate system through time alignment and coordinate transformation, establishing an initial spatial association for the point clouds across consecutive frames. The specific process includes data parsing and indexing, cross-frequency data time alignment, and coordinate transformation and initial registration.
[0045] The data parsing and indexing process is as follows: First, the data packets recorded by the ground truth acquisition vehicle are split and parsed to extract two types of core data (i.e., LiDAR point cloud data and GNSS / IMU pose data) and an index file is created. Specifically, for the LiDAR point cloud data: the raw LiDAR point cloud is converted into a Point Cloud Data (PCD) file format using a Point Cloud Library (PCL). Each frame of the point cloud contains XYZ (relative to vehicle body coordinates) coordinates, intensity, timestamp, etc., and a pcd_timestamp.txt index file is generated, recording the index value, acquisition timestamp, and storage path of each frame. For the GNSS / IMU pose data: core information such as position (x, y, z) and attitude (quaternions qx, qy, qz, qw) output by the positioning system is extracted, and an odometry_loc.txt index file is generated, containing the pose data index, timestamp, coordinate values, and variance information. Its positioning accuracy can reach within 10cm, providing a high-precision benchmark for registration.
[0046] The cross-frequency data time alignment process is as follows: The acquisition frequencies of lidar and GNSS / IMU differ (e.g., lidar acquisition frequency is 10Hz, GNSS / IMU acquisition frequency is 100Hz), resulting in no direct overlap of timestamps between the two types of data. Accurate alignment requires time interpolation. Specifically, the alignment process involves using the timestamp of the point cloud frame as a reference, finding the two closest valid pose data points before and after that timestamp in the high-frequency pose data of the GNSS / IMU. A linear interpolation algorithm is then used based on the time difference to calculate the pose information (including position and attitude) that perfectly matches the acquisition time of the current point cloud frame, ensuring that each point cloud frame can be matched with a unique, high-precision pose reference.
[0047] The coordinate transformation and initial registration process is as follows: Using the aligned pose information, each frame of the point cloud is transformed from the vehicle-relative coordinate system of the LiDAR to a global coordinate system (such as WGS84), completing the initial registration of the continuous frame point clouds. The initial registration process involves: calculating the rotation matrix R using pose quaternions to convert the relative pose of the point cloud into the global pose; combining the position coordinates to calculate the translation vector t, forming a complete rigid transformation matrix T. All points (x, y, z) in each frame of the point cloud are converted into global coordinates (X, Y, Z) using the transformation matrix T, enabling the continuous frame point clouds to establish spatial association based on their real physical locations, avoiding point cloud misalignment caused by vehicle movement, and providing reliable initial values for subsequent fine registration.
[0048] Step 3: Optimize the relative poses between point clouds using algorithms such as Iterative Closest Point (ICP) to improve the overall map accuracy.
[0049] Even after initial registration via GNSS / IMU, errors still exist between point clouds. Therefore, precise registration is required. The core function of the ICP algorithm is to correct these deviations through a concise "matching-optimization-iteration" process to achieve accurate point cloud registration. The specific process of optimizing the relative pose between point clouds using the ICP algorithm is as follows: First, select two consecutive frames of point clouds (e.g., the point cloud after initial registration in the current frame and the point cloud registered in the previous frame). The point cloud registered in the previous frame is designated as the "target point cloud," and the point cloud after initial registration in the current frame is designated as the "source point cloud." The goal of the ICP algorithm is to find the optimal rotation + translation transformation relationship so that the source point cloud, after transformation, accurately coincides with the target point cloud. To improve efficiency, voxel filtering is first used to reduce the number of point clouds.
[0050] Then, optimization iterations are performed. The core iterative process is as follows: using an efficient KD-tree algorithm, for each point in the source point cloud, the nearest point in the target point cloud is found, forming a "point pair". Simultaneously, outlier point pairs that are too far apart are filtered out to avoid interference. Based on the filtered point pairs, the rotation matrix and translation vector that minimize the sum of distances between these point pairs are calculated using mathematical methods (singular value decomposition), which is the current optimal transformation relationship. The source point cloud position is updated using the calculated transformation relationship, and then the average distance error between the updated source point cloud and the target point cloud is calculated. If the average distance error is small enough to meet the requirements (e.g., less than 0.01m), or the number of iterations reaches the preset number of iterations (e.g., 50 times), the iteration stops; otherwise, the process returns to selecting two consecutive frames of point clouds, using the updated point cloud to find the corresponding points again and calculate the transformation relationship, iterating repeatedly until the conditions are met (e.g., reaching the preset number of iterations).
[0051] Step 4: Stitch and merge the registered point clouds from multiple frames to generate a dense 3D point cloud map covering the target parking scene (e.g., underground garage, open-air parking lot, park road).
[0052] Among them, the 3D point cloud map can clearly present the geometric outlines of static targets such as parking lines, curbs, pillars, and walls.
[0053] Specifically, the process of stitching and fusing multiple registered point clouds to generate a dense 3D point cloud map covering the target parking scene mainly includes: global coordinate system I and point cloud projection, global coordinate system I and point cloud projection, sparse region completion and detail enhancement, and sparse region completion and detail enhancement.
[0054] The global coordinate system and point cloud projection process are as follows: using the global coordinates of the first frame of the point cloud as a reference, all subsequent frame point clouds that have undergone precise registration are projected into the same global coordinate system through their respective final transformation matrices T. At this point, the spatial positions of all point clouds accurately correspond to the real physical environment, laying the foundation for stitching.
[0055] The process of deduplication and redundant point filtering in overlapping areas is as follows: During vehicle movement data acquisition, there are many overlapping areas in the point clouds of adjacent frames (e.g., the same pillar or parking line), which need to be deduplicated through spatial clustering. The spatial clustering deduplication process is as follows: A voxel grid filtering algorithm is used to divide the global space into small voxels (e.g., 5cm×5cm×5cm), and only one representative point (e.g., the center point or the point closest to the centroid) is retained in each voxel. Then, point cloud intensity information is used to assist in deduplication. Specifically, for points in overlapping areas, points with higher intensity values are preferentially retained (usually corresponding to clearer outlines of object surfaces) to improve map detail accuracy.
[0056] The sparse region completion and detail enhancement process is as follows: For sparse regions caused by blind spots in LiDAR scanning (such as parking space corners and pillar shadows), a neighborhood interpolation-based completion algorithm is used to identify holes in the sparse regions. Using the coordinates and normal vectors of surrounding dense points as constraints, completion points conforming to physical laws are generated. Detailed information from multiple frame point clouds is fused; for example, if a curb is not clearly scanned in a particular frame, it is supplemented using overlapping points from subsequent frames to ensure the integrity of key targets such as parking lines and pillar outlines.
[0057] The map post-processing and verification process involves: globally smoothing the stitched point cloud to remove a small number of residual outliers (such as floating points in the air or noise points); and verifying the map using visualization tools (such as PCL Visualizer) to check for obvious misalignments, holes, or other issues, ensuring that the map clearly presents the geometric outlines of all static targets in the target parking scene.
[0058] Point cloud annotation is performed using a semi-automated annotation tool based on the generated 3D point cloud map. Annotated targets include static targets, dynamic targets, and drivable areas (Freespace). Static targets include, for example, parking spaces (annotating their corner points or bounding boxes) and static obstacles (such as pillars, annotating their 3D bounding boxes). Dynamic targets include, for example, pedestrians and vehicles (annotating their 3D bounding boxes and trajectories). Drivable areas are defined as the outlines of areas where safe driving is permitted. The annotation results generate structured annotation files (e.g., JSON format) containing the category label, 3D coordinates in the global coordinate system, size, orientation angle, and other attribute information for each target. These annotation files serve as high-quality ground truth data for the initial training and testing of the target perception model.
[0059] The semi-automated annotation tool includes both automated and interactive functions. The automated function supports automatic identification of targets such as pillars and vehicles based on point cloud clustering algorithms and generates initial 3D bounding boxes. By fusing with image data, it improves the annotation accuracy of static targets (such as parking space corners). The interactive function provides features such as bounding box drag-and-drop adjustment, corner fine-tuning, and quick category switching. It supports automatic polygon fitting and manual correction for freespace areas, and annotation results can be directly exported as JSON format, containing complete attributes such as global coordinates, dimensions, and orientation angle of the target.
[0060] In some embodiments, after acquiring the target image data, the method further includes: annotating and projecting the target image data according to a preset annotation tool and a preset inverse perspective transformation algorithm to obtain a bird's-eye view planar coordinate representation of the image data in the vehicle coordinate system; and generating a target format annotation file based on the image data in the bird's-eye view planar coordinate representation.
[0061] The preset annotation tool is a semi-automated annotation tool, the preset inverse perspective transformation algorithm is based on the principle of inverse perspective transformation, and the target format is the standard JSON format.
[0062] 2D image data supplementation is the core of multi-task BEV perception models, enabling data supplementation and rapid iteration at any stage (e.g., testing, mass production, actual operation) where problems are identified and supplementary data is needed. It is used for efficient data collection and annotation of specific problem scenarios (such as irregularly shaped parking spaces, corners with temporary piles of debris, and curbs made of special materials) that are discovered during testing but not covered in the initial dataset.
[0063] Figure 6 This is a schematic diagram of the 2D data supplementation process provided in the embodiments of this application. For example, see [link to relevant documentation]. Figure 6 The 2D image data supplement includes: 2D image data acquisition, 2D image data annotation and projection, and generation of annotation files in the target format. Specifically, 2D image data acquisition utilizes mass-produced automotive surround-view fisheye cameras or front-view cameras to capture image data of the problem scene. The acquisition process must ensure coverage of diverse environmental conditions, including different lighting conditions (such as strong light, backlight, and low light), different weather conditions (such as sunny days and cloudy days), and different scene complexities (such as single obstacles and densely packed multiple obstacles), to avoid data bias in the model and improve its generalization ability.
[0064] 2D image data annotation and projection involves manually or semi-automatically annotating the acquired 2D image data to mark the outlines or polygon vertices of static targets (such as parking spaces and obstacles). Subsequently, using known camera intrinsic parameters (such as focal length and principal point coordinates) and extrinsic parameters (camera pose relative to the vehicle body), the coordinates of the polygon vertices annotated on the image plane are projected onto the BEV plane (usually the ground plane) in the vehicle coordinate system through the principle of Inverse Perspective Mapping (IPM). This converts the 2D image annotations into coordinate representations in BEV space, aligning them with the output format of the multi-task BEV perception model and facilitating subsequent training of the model.
[0065] Finally, the above annotation files are used to generate a final standard JSON format annotation file. This file contains the image storage path, the category information of each detected target, and its polygon vertex coordinate sequence on the BEV plane. These files are merged with the initial point cloud annotation dataset and used together for incremental training of the multi-task BEV perception model. For example, using a programming language (such as Python), the organized information is generated into JSON format text according to preset field rules. For example, the entire file content is enclosed in curly braces {}, and each field is represented by a key-value pair "key": value (strings are enclosed in double quotes, numbers are written directly, and arrays are enclosed in square brackets []). Multiple targets are enclosed in arrays [], and each target is represented by curly braces {} to indicate its independent information. For example, suppose the annotation result of an image is: Image path: . / data / images / 001.jpg Target 1: Category "Parking Space", BEV coordinates (clockwise): (1.2, 0.8), (3.8, 0.8), (3.8, 2.2), (1.2, 2.2) Target 2: Category "Obstacle", BEV coordinates (clockwise): (5.1, 3.3), (5.7, 3.3), (5.7, 4.1), (5.1, 4.1).
[0066] Step 120: Train and optimize the target perception model based on the image data to obtain the trained target perception model.
[0067] Among them, the trained target perception model is used to perceive multiple target tasks for automatic parking.
[0068] In some embodiments, the target perception model includes a backbone network, a feature pyramid network, a first detection head, a second detection head, and a third detection head. Training and optimizing the target perception model based on image data to obtain a trained target perception model includes: extracting features from target image data using the backbone network to obtain a first feature map; wherein the first feature map includes semantic features and positional features; fusing the semantic features and positional features using a feature pyramid network (FPN) to obtain a second feature map; and training and optimizing the first detection head, the second detection head, and the third detection head based on the second feature map to obtain a trained target perception model; wherein multiple target tasks include a first target task, a second target task, and a third target task; the first detection head is used to output the first target task, the second detection head is used to output the second target task, and the third detection head is used to output the third target task.
[0069] The backbone network employs the ResNet series, such as ResNet-50, to extract deep features from the target image data. The target image data consists of single or multiple frames captured by a standard vehicle-mounted camera. The second feature map is a multi-scale feature map.
[0070] Specifically, the feature extraction process for the target image data is as follows: Single or multiple frames of image data (i.e., target image data) captured by the vehicle-mounted camera are input into the backbone network. The backbone network extracts deep features from the target image data, resulting in the first feature map. To fuse semantic information at different scales, a Feature Pyramid Network (FPN) is introduced to fuse the semantic and positional features of the first feature map. Specifically, the FPN fuses strong semantic features from higher levels with fine positional features from lower levels through top-down paths and lateral connections, outputting a set of multi-scale feature maps (P2, P3, P4, P5).
[0071] After feature extraction from the target image data, the process also includes transforming the second feature map into the BEV (Bird's Eye View) coordinate system. Specifically, inverse perspective transformation (IPM) is used to transform the image feature map (i.e., the second feature map) from the image coordinate system to the BEV coordinate system. Specifically, a height is preset for each point on the second feature map, and then, using camera calibration parameters (intrinsic and extrinsic parameters) and the IPM transformation matrix, the image feature points are projected onto their corresponding positions on the BEV grid, forming the BEV feature map. The advantage of this transformation is its computational efficiency and suitability for real-time applications.
[0072] The training method for the perception model provided in this application can train and optimize a multi-task BEV perception model, enabling it to perceive multiple target tasks in a vehicle parking scenario. These multiple target tasks include a first target task, a second target task, and a third target task. The multi-task BEV perception model includes a dedicated multi-task detection head. This head comprises three structurally independent and functionally specialized detection heads: a first detection head, a second detection head, and a third detection head. Each detection head performs a different task. The first detection head (i.e., the static target detection head) outputs the first target task, the second detection head (i.e., the dynamic target detection head) outputs the second target task, and the third detection head (i.e., the drivable area extraction head (Freespace head)) outputs the third target task. The first target task is to detect and classify static targets. The second target task is to detect dynamic targets and predict their motion state. The third target task is to identify the safe drivable area.
[0073] In some embodiments, the first objective task is to detect and classify static targets; the first detection head includes a convolutional network, a first output head, and a second output head; the method for training and optimizing the first detection head based on a second feature map includes: inputting the second feature map into the convolutional network, and outputting a target class probability map and a bounding box regression map through the first output head and the second output head, respectively; optimizing the parameters of the target class probability map and the bounding box regression map according to a first preset loss function, so as to detect and classify static targets through the optimized first detection head.
[0074] The first detection head is used to detect and classify static targets. Static targets include standard parking spaces, irregularly shaped parking spaces, pillars, curbs, etc.
[0075] The network structure of the first detection head includes a convolutional network, a first output head, and a second output head. The convolutional network is a small convolutional network, typically consisting of four 3x3 convolutional layers (with the number of channels decreasing layer by layer, such as 256, 128, 64, 32). The number of channels equals the number of static target categories (e.g., three categories: parking space, pillar, curb). The first and second output heads are two independent 1x1 convolutional layers.
[0076] The first preset loss function is a joint loss function. Specifically, the loss for the target class probability map (i.e., the classification task) uses Focal Loss to address the foreground-background class imbalance. The loss for the bounding box regression map (i.e., the regression task) uses Generalized Intersection over Union (GIoU) Loss to better optimize the overlap of bounding boxes. The total loss is a weighted sum of these two losses; for example, the ratio of Focal Loss weights to GIoU Loss weights is 1:2.
[0077] Specifically, the training and optimization process for the first detection head is as follows: After acquiring the shared BEV feature map (i.e., the second feature map), the shared BEV feature map is input into a convolutional network. The convolutional network extracts the target class probability map and bounding box regression map from the shared BEV feature map, and outputs the target class probability map and bounding box regression map through the first and second output heads, respectively. The target class probability map is used to detect and classify the number of static target categories, and the bounding box regression map is used to detect the bounding boxes of static targets, predicting the bounding box parameters for each possible location of a static target, for example, using a center-based representation (predicting the distance from the center point to the left, right, front, and back edges of the bounding box). Finally, the parameters of the target class probability map and bounding box regression map are optimized according to a first preset loss function to obtain the optimized first detection head, which is used to detect and classify static targets. Therefore, by sharing the BEV feature map and the first detection head, the training and optimization of the detection and classification capabilities of static targets can be achieved, thereby realizing the training and optimization of the multi-task BEV perception model for the detection and classification of static targets, and thus improving the accuracy and reliability of static target perception in subsequent automatic parking.
[0078] In some embodiments, the second objective task is to detect a dynamic target and predict its motion state; the second detection head includes a first encoder; the method for training and optimizing the second detection head based on the second feature map includes: inputting a second feature map of a preset number of consecutive frames into the first encoder; capturing the correlation between dynamic targets in adjacent frames through the first encoder to track the dynamic target and estimate its velocity; predicting the interval to which the dynamic target belongs and its offset relative to the midpoint of the interval based on a preset binning regression algorithm; determining the heading angle of the dynamic target based on the interval and the offset; and optimizing the velocity and heading angle parameters according to a second preset loss function to detect the dynamic target and predict its motion state using the optimized second detection head.
[0079] The second detection head is responsible for detecting dynamic targets and predicting their motion state. Dynamic targets include pedestrians and vehicles. Motion state includes the target's speed and heading angle.
[0080] The network structure of the second detection head includes: a first encoder. The first encoder is a Transformer encoder.
[0081] The preset number of frames is N, where N is an integer, such as 5 frames. The specific value can be set according to the actual situation, and no specific limit is set here.
[0082] The specific process of tracking and estimating the velocity of a dynamic target is as follows: To incorporate temporal information, the BEV feature maps of N consecutive frames (e.g., 5 frames) are concatenated along the channel dimension or input into a Transformer encoder. The Transformer utilizes a self-attention mechanism to capture the correlation between targets across frames, thereby tracking and estimating the velocity of the dynamic target. Specifically, the implementation process of using the self-attention mechanism to capture the correlation between targets across frames, thereby tracking and estimating the velocity of the dynamic target, includes the following steps: Step 1: Input Sequence Construction. Flatten the BEV feature maps of N consecutive frames into a sequence, with each frame having a feature shape of [H×W, C], and stack them as a temporal input [N, H×W, C]. Add positional encoding (spatial + temporal) to distinguish different frames and pixel locations.
[0083] Step 2: Self-Attention Calculation. Utilizing a self-attention mechanism, the system automatically finds the corresponding positions of the same target in different frames by comparing the feature similarity of various locations in the BEV feature map across different frames, thereby establishing the target's motion trajectory and calculating its velocity. Specifically, this includes: 1. Feature Mapping (QKV Generation). For the BEV feature map, three sets of features are calculated: Query, Key, and Value. Query represents the positions in the current frame, used to actively find matches. Key represents the positions in other frames, used for matching. Value contains detailed information about the target, used for subsequent calculations. 2. Calculating Inter-Frame Associations (Attention Weights). The feature similarity of each position (Query) in the current frame is compared with all positions (Key) in other frames. Position pairs with high similarity (such as the vehicle's position in adjacent frames) receive high weights. Softmax normalization concentrates the weights on truly matching targets, filtering out irrelevant regions. 3. Feature Aggregation (Motion Information Extraction). Based on the calculated weights, the Value features of other frames are weighted and fused to ensure that the features of the same target remain consistent across different frames. High-weighted matching pairs naturally imply the target's displacement information (e.g., the vehicle moved 2 meters from the previous frame). IV. Multi-head attention (improved robustness). Multiple independent QKV calculations (multi-head) are used to focus on different characteristics of the target (e.g., position, shape, motion trend), and the results are finally merged to make tracking more stable.
[0084] Step 3: Determine velocity and trajectory. Displacement calculation: The displacement (Δx, Δy) of each target is decoded from the aggregated features using a regression network. Velocity estimation: Integrating the time interval Δt, the instantaneous velocity is obtained by bit removal. Target tracking: High-weighted matching of the same target across different frames automatically forms the trajectory, eliminating the need for additional data association algorithms.
[0085] The process of determining the heading angle of a dynamic target is as follows: To address the problem of inaccurate heading angle prediction for dynamic targets, a "binning + regression" strategy (i.e., a pre-set binning and regression algorithm) is adopted. The 360-degree range is divided into K (e.g., 10) angle intervals (bins). The network first predicts which interval the dynamic target belongs to (classification problem), then predicts a small offset relative to the midpoint of the interval (regression problem), and finally combines binning and regression to obtain the accurate heading angle. The accurate heading angle = the midpoint angle of the interval to which the target belongs + the small offset predicted by the network regression. For example, assuming K=10 (interval width 36°), the classification predicts the dynamic target belongs to interval 4 (midpoint 126°), and the regression predicts an offset of +2.3°, then the accurate heading angle = 126° + 2.3° = 128.3°.
[0086] The second preset loss function employs a multi-task loss based on the Hungarian algorithm for matching. First, the Hungarian algorithm is used to optimally match the predicted bounding boxes with the ground truth bounding boxes. For each matched target pair, the L1 loss for its center point coordinates, the class cross-entropy loss, the MSE loss for velocity, and the heading angle loss are calculated. These losses are then summed according to a certain weight (e.g., 2:1:1:1) to obtain the total loss.
[0087] In this model, the predicted bounding box (BWT) refers to the rectangular bounding box predicted by the model for regions in an image where objects may exist. It typically includes attributes such as center coordinates (x / y), length and width (w / l), and class probability. The ground truth bounding box (True Box) refers to the actual location and class of the object in the image, with the same format as the predicted bounding box. A matched object pair is the combination obtained by matching a predicted bounding box with a ground truth bounding box one-to-one using the Hungarian algorithm. Each ground truth bounding box matches at most one predicted bounding box, and vice versa. The matching principle is cost minimization: typically based on the joint cost of IoU (Intersection over Union) or classification error. For example, predicted bounding box 1... Ground truth bounding box 1 (IoU=0.85, class consistent), predicted bounding box 2 The ground truth bounding box 2 (IoU = 0.80, class consistent) is considered. Unmatched predicted bounding boxes (such as low-confidence boxes) are treated as negative samples or redundant detections. Specifically, the process of using the Hungarian algorithm to optimally match predicted and ground truth bounding boxes is as follows: Constructing a cost matrix: Calculating the cost between all predicted and ground truth bounding boxes (e.g., 1 - IoU + classification error). Solving with the Hungarian algorithm: Finding the unique matching combination with the minimum total cost. For example, referring to the simplified cost matrix example in Table 1, the optimal solution is: Predicted bounding box 1 matches ground truth bounding box 1 (cost 0.15), predicted bounding box 2 matches ground truth bounding box 2 (cost 0.20), total cost = 0.35 (globally minimum).
[0088] Table 1: Example Table of Cost Matrix
[0089] Specifically, the training and optimization process for the second detection head is as follows: After acquiring the shared BEV feature map (i.e., the second feature map), the shared BEV feature map of N consecutive frames (e.g., 5 frames) is input into the first encoder. The first encoder uses a self-attention mechanism to capture the correlation between targets between frames, thereby tracking dynamic targets and estimating their speed. A preset binning regression algorithm is used to predict the interval to which the dynamic target belongs and its offset relative to the midpoint of the interval. The heading angle of the dynamic target is determined based on the interval and the offset. The speed and heading angle are then optimized using a second preset loss function to obtain the optimized second detection head. This optimized second detection head is used to detect dynamic targets and predict their motion state. Thus, by using the shared BEV feature map and the second detection head, the training and optimization of the ability to detect and predict the motion state of dynamic targets is achieved. This enables the multi-task BEV perception model to train and optimize its ability to detect and predict dynamic targets, thereby improving the accuracy and reliability of dynamic target perception in subsequent automatic parking.
[0090] In some embodiments, the third objective task is to identify a safe driving area; the third detection head includes a second encoder and a first decoder; the method for training and optimizing the third detection head based on the second feature map includes: outputting a target segmentation map based on the second encoder, the first decoder, and the second feature map; outputting a binary segmentation map of the safe driving area and an edge map of obstacles based on the target segmentation map and a preset line-surface joint learning algorithm; and optimizing the parameters of the binary segmentation map and the edge map based on a third preset loss function so as to identify the safe driving area through the optimized third detection head.
[0091] The third detection head is used to identify safe driving areas through pixel-level segmentation. The network structure of the third detection head adopts a variant of the encoder-decoder structure (i.e., including a second encoder and a first decoder). The encoder part (i.e., the second encoder) shares BEV features with the static target detection head (i.e., the first detection head), and the decoder part (i.e., the first decoder) is usually composed of several transposed convolutional or upsampling layers to progressively restore the resolution, and finally outputs a target segmentation map with the same spatial resolution as the input BEV grid.
[0092] The preset line-surface joint learning algorithm is a line + surface joint learning algorithm. To improve edge accuracy, in addition to the main segmentation task (outputting a binary segmentation map of the freespace region, i.e., "surface"), an auxiliary task branch (i.e., edge prediction task) is added to predict the edge map of obstacles (i.e., "lines"). During training, the two tasks (i.e., the main segmentation task and the edge prediction task) are jointly supervised. During inference, edge information can be used to post-process the segmentation results. For example, edge optimization can be performed using a Conditional Random Field (CRF).
[0093] The third preset loss function is as follows: Dice Loss is used for the main segmentation task, which can effectively handle the area imbalance problem between the foreground and background. Binary Cross-Entropy Loss (BCE Loss) is used for the edge prediction task. The total loss is the weighted sum of the two (for example, the ratio of Dice Loss weight to BCE Loss weight is 1:2).
[0094] The drivable area extraction head is based on an encoder-decoder architecture. It extracts abstract features through multi-level downsampling, then restores spatial resolution through progressive upsampling, ultimately outputting a pixel-level segmentation map with the same size as the input BEV mesh. Efficient information fusion between the encoder and decoder is achieved through feature sharing and skip connections. Specifically: Encoder (downsampling): Based on shared BEV features, a second encoder (such as an additional convolutional layer) reduces the feature map resolution, enhancing semantic representation capabilities. Decoder (upsampling): Resolution is gradually restored through transposed convolution or interpolation upsampling (such as bilinear interpolation), and skip connections inject intermediate layer features from the encoder to optimize edge details.
[0095] Specifically, the training and optimization process for the third detection head is as follows: After acquiring the shared BEV feature map (i.e., the second feature map), the shared BEV feature map is input into the second encoder. The second encoder and the first encoder then output a target segmentation map with the same grid spatial resolution as the shared BEV feature map. Next, based on the target segmentation map and a preset line-surface joint learning algorithm, a binary segmentation map of the safe driving area and edge maps of obstacles are output. Finally, the parameters of the binary segmentation map and edge maps are optimized according to a third preset loss function to identify the safe driving area using the optimized third detection head. Thus, by using the shared BEV feature map and the third detection head, the training and optimization of the segmentation and recognition capability of the safe driving area are achieved. This enables the training and optimization of the multi-task BEV perception model for the safe driving area recognition task, thereby improving the accuracy and reliability of safe driving area recognition in subsequent automatic parking.
[0096] Figure 7 This is a schematic diagram illustrating the training process of the multi-task BEV perception model provided in this embodiment. For an example, please refer to [link to example]. Figure 7 The training process of this multi-task BEV perception model includes: inputting single-frame or multi-frame images from an onboard camera (i.e., target image data); extracting image features from the target image data using ResNet+FPN to obtain a multi-scale feature map (i.e., the first feature map); then, performing BEV space transformation on the multi-scale feature map using the IPM transformation principle to obtain a BEV feature map (i.e., the second feature map); outputting the BEV feature map to a dedicated multi-task detection head as a shared BEV feature map for the first, second, and third detection heads; finally, each detection head is trained and optimized based on the shared BEV feature map and its own network structure to obtain the trained and optimized first, second, and third detection heads, thus obtaining the trained and optimized multi-task BEV perception model. This improves the accuracy and reliability of perception of static targets, dynamic targets, and safe driving areas in automatic parking. Therefore, by using a single forward-looking or surround-view camera input, the three core perception tasks in parking scenarios—static target detection, dynamic target detection, and freespace segmentation—can be simultaneously achieved in BEV space. Furthermore, the target perception model adopts an architecture of "shared backbone network + multi-task dedicated head", which can significantly improve computing efficiency while ensuring high accuracy, thereby meeting the real-time requirements of the vehicle embedded platform.
[0097] Figure 8 This is an overall flowchart of the multi-task BEV perception model training method provided in the embodiments of this application. (See also...) Figure 8 The overall process includes: First, dataset construction and partitioning. For the initial dataset: base ground truth data generated using point cloud map construction and annotation methods, such as an initial dataset containing 500,000 frames of labeled data. For the incremental dataset: approximately 50,000 frames of data collected and labeled using 2D image data supplementation methods to address new problem scenarios discovered during mass production. After merging the initial and incremental datasets, they are randomly partitioned into training, validation, and test sets in a 7:2:1 ratio.
[0098] Next, data augmentation: Online data augmentation techniques are applied to the images in the training set to improve the robustness of the multi-task BEV perception model. These online data augmentation techniques include: Geometric transformations: random horizontal flipping (probability 50%), random cropping (scale 0.8-1.0); and photometric transformations: random adjustment of brightness (±20%), contrast, addition of Gaussian noise or blur (standard deviation 0-0.5).
[0099] Secondly, the training parameters are configured as follows: First, the optimizer is configured: the AdamW optimizer is used, with an initial learning rate of 1e-3 and a weight decay coefficient of 1e-4 to control overfitting. Second, the learning rate scheduling is configured: a cosine annealing strategy is used, causing the learning rate to decrease from its initial value to its minimum value (e.g., 1e-6) according to a cosine curve during training. Third, the training settings are: the batch size is set to 16, and the total training duration is 40 epochs. The loss is monitored on the validation set; if the loss no longer decreases for 5 consecutive epochs, early stopping is triggered to prevent overfitting and restore the model parameters that best perform on the validation set.
[0100] In summary, the training method for the perception model provided in this application constructs a complete technical system from high-precision data production and dedicated model design to agile data iteration. It has the following advantages: First, a high-precision ground truth automatic generation method based on point cloud maps is provided. Specifically, a precise 3D map of the parking scene is constructed using LiDAR point clouds as a "benchmark," and targets are labeled on this basis. This method fundamentally avoids the geometric distortion problems caused by perspective and occlusion in traditional 2D image labeling, providing a reliable "ground truth" data source for target perception model training.
[0101] Secondly, a decoupled multi-task BEV perception model architecture optimized specifically for automated parking scenarios is provided. Specifically, it employs a "shared feature encoding + task-specific decoding head" design, rather than a simple multi-task concatenation. Instead, customized network substructures and loss functions are designed for each of the three core tasks: static object detection, dynamic object tracking, and drivable area segmentation. For example, the static object detection head uses a weighted combination of Focal Loss and GIoU Loss to address class imbalance and improve bounding box accuracy; the dynamic object detection head introduces a temporal attention mechanism (such as a Transformer encoder) and a heading angle binning regression strategy to achieve accurate tracking; and the Freespace head uses a joint learning strategy of "surface segmentation + edge detection" to optimize contour details.
[0102] Third, it provides a low-cost, agile data closed-loop iteration mechanism for mass production. Specifically, it establishes a dual-track data pipeline of "high-precision ground truth initialization of point clouds and rapid supplementation of 2D images". In the mass production stage, for new scenarios, there is no need to use expensive ground truth acquisition vehicles. Only 2D images need to be acquired using the cameras of the mass production vehicle. After annotation, they can be projected onto the BEV space through inverse perspective transformation to generate effective training data.
[0103] Fourth, this application provides an automatic parking perception model training system that integrates the above methods. The system includes a point cloud map construction and annotation module, a 2D image data supplementation module, a multi-task BEV perception model training module, and the collaborative working relationships and data flow logic between these modules, collectively achieving high-quality and high-efficiency model training and iteration.
[0104] Furthermore, compared with related technologies, the application specifically claims the following advantages: First, at the data foundation level, this application employs a labeling method based on LiDAR point cloud maps. Point cloud data itself possesses precise three-dimensional geometric information and is unaffected by 2D image perspective and occlusion. Therefore, it can be logically deduced that this application can significantly improve the geometric accuracy of the training data's true values from the source, laying a reliable data foundation for addressing the problem of insufficient accuracy in perception models.
[0105] Secondly, regarding model performance, this application employs a multi-task BEV perception model architecture and loss function (such as GIoU loss for static heads, temporal attention for dynamic heads, and edge optimization for Freespace heads) that is deeply customized for the characteristics of different perception tasks. This decoupled and specialized design allows each task to achieve its optimal learning objective within a shared feature context. Therefore, it can be logically deduced that this application can simultaneously and significantly improve static target detection accuracy, dynamic target tracking stability, and drivable region segmentation edge details, effectively solving the problem of difficult collaborative optimization of multi-task performance in general models.
[0106] Thirdly, regarding engineering implementation and mass production efficiency, this application establishes an agile iterative mechanism that supplements BEV data with rapid projection of 2D images. This mechanism avoids absolute dependence on high-cost ground truth acquisition vehicles. Therefore, it can be inferred that this application can significantly reduce the cost and time cycle of acquiring data for new scenarios, enabling the model to quickly respond to mass production needs and continuously iterate and optimize, fundamentally solving the pain point of lagging model iteration.
[0107] In summary, this application, through interconnected technological innovations, has brought substantial progress in three dimensions: data quality, model performance, and mass production efficiency, forming a complete and superior solution.
[0108] Figure 9 This is a block diagram illustrating the principle structure of a training system for a perception model provided in an embodiment of this application. This application also provides a training system for a perception model, see below. Figure 9The training system 100 for the perception model includes: a first acquisition module 101 for acquiring target image data; wherein the target image data is two-dimensional image data of a vehicle in a target parking scene acquired by an image acquisition module; a second acquisition module 102 for acquiring a target perception model; the target perception model is a multi-task bird's-eye view perception model; and a training module 103 for training and optimizing the target perception model based on the target image data to obtain a trained target perception model; wherein the trained target perception model is used to perceive multiple target tasks.
[0109] The technical solution of this application provides a training system for a perception model. This application constructs a perception model based on a multi-task bird's-eye view and uses two-dimensional image data of the vehicle in the target parking scene acquired by the image acquisition module to train and optimize the multi-task perception model, thereby achieving efficient collaborative training of multi-task perception for automatic parking.
[0110] In some embodiments, the target perception model includes a backbone network, a feature pyramid network, a first detection head, a second detection head, and a third detection head; the training module 103 is further configured to: The target image data is feature extracted based on the backbone network to obtain the first feature map; wherein the first feature map includes semantic features and positional features; The semantic and positional features are fused using a feature pyramid network to obtain a second feature map. The first, second, and third detection heads are then trained and optimized based on the second feature map to obtain a trained target perception model. The model includes multiple target tasks, such as a first target task, a second target task, and a third target task. The first detection head is used to output the first target task, the second detection head is used to output the second target task, and the third detection head is used to output the third target task.
[0111] In some embodiments, the first objective task is to detect and classify static targets; the first detection head includes a convolutional network, a first output head, and a second output head; the training module 103 is further configured to: input the second feature map into the convolutional network, and output the target category probability map and the bounding box regression map through the first output head and the second output head, respectively; The parameters of the target category probability map and bounding box regression map are optimized according to the first preset loss function so that the static target can be detected and classified by the optimized first detection head.
[0112] In some embodiments, the second objective task is to detect a dynamic target and predict the motion state of the dynamic target; the second detection head includes a first encoder; the training module 103 is further configured to: The second feature map of a preset number of consecutive frames is input into the first encoder; The first encoder captures the correlation between dynamic targets in adjacent frames to track and estimate the speed of the dynamic targets. The dynamic target is predicted to belong to the interval and its offset relative to the midpoint of the interval based on the preset binning regression algorithm. The heading angle of the dynamic target is determined based on the interval and the offset. The velocity and heading angle are optimized according to the second preset loss function so that the dynamic target can be detected by the optimized second detection head and the motion state of the dynamic target can be predicted.
[0113] In some embodiments, the third objective task is to identify a safe driving area; the third detection head includes a second encoder and a first decoder; the training module 103 is further configured to: The target segmentation map is output based on the second encoder, the first decoder, and the second feature map. Based on the target segmentation map and the preset line-surface joint learning algorithm, the binary segmentation map of the safe driving area and the edge map of the obstacle are output. The parameters of the binary segmentation map and edge map are optimized according to the third preset loss function so that the safe driving area can be identified by the optimized third detection head.
[0114] In some embodiments, the training system 100 for the perception model further includes: The annotation and projection module is used to annotate and project the target image data according to the preset annotation tools and preset inverse perspective transformation algorithms, so as to obtain the bird's-eye view planar coordinate representation of the target image data in the vehicle coordinate system. The generation module is used to generate a target format annotation file based on the target image data represented by the planar coordinates of the bird's-eye view.
[0115] In some embodiments, the vehicle further includes a point cloud acquisition module; the training system 100 for the perception model also includes: The training optimization module is used to train and optimize the target perception model in the early stages of training by acquiring point cloud data collected by the point cloud acquisition module.
[0116] This embodiment also provides a method including a memory and a processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, it implements any of the methods described in the above embodiments.
[0117] This embodiment also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of any of the methods in the above embodiments.
[0118] In the embodiments of this application, the storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.
[0119] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0120] The foregoing has provided a detailed description of a training method, system, device, and storage medium for a perception model provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for training a perception model, the method comprising: An automatic parking system for vehicles, wherein the vehicle includes an image acquisition module; the training method includes: Acquire target image data and target perception model; wherein, the target image data is two-dimensional image data of the vehicle in the target parking scene acquired by the image acquisition module; the target perception model is a pre-constructed multi-task bird's-eye view perception model; The target perception model is trained and optimized based on the image data to obtain a trained target perception model; wherein the trained target perception model is used to perceive multiple target tasks of the automatic parking.
2. The method of claim 1, wherein, The target perception model includes a backbone network, a feature pyramid network, a first detection head, a second detection head, and a third detection head; the step of training and optimizing the target perception model based on the target image data to obtain a trained target perception model includes: The target image data is subjected to feature extraction based on the backbone network to obtain a first feature map; wherein, the first feature map includes semantic features and positional features; The semantic features and the positional features are fused according to the feature pyramid network to obtain a second feature map; the first detection head, the second detection head, and the third detection head are trained and optimized according to the second feature map to obtain the trained target perception model; wherein, the plurality of target tasks include a first target task, a second target task, and a third target task; the first detection head is used to output the first target task, the second detection head is used to output the second target task, and the third detection head is used to output the third target task.
3. The method of claim 2, wherein, The first objective is to detect and classify static targets; the first detection head includes a convolutional network, a first output head, and a second output head. A method for training and optimizing the first detection head based on the second feature map includes: The second feature map is input into the convolutional network, and the target class probability map and the bounding box regression map are output through the first output head and the second output head, respectively. The target category probability map and the bounding box regression map are optimized according to the first preset loss function so that the static target can be detected and classified by the optimized first detection head.
4. The method of claim 2, wherein, The second objective task is to detect dynamic targets and predict the motion state of the dynamic targets; the second detection head includes a first encoder; the method for training and optimizing the second detection head based on the second feature map includes: The second feature map of a preset number of consecutive frames is input into the first encoder; The first encoder captures the correlation between the dynamic targets in adjacent frames to track the dynamic targets and estimate their speed. The dynamic target is predicted to belong to the interval and its offset relative to the midpoint of the interval based on a preset binning regression algorithm. The heading angle of the dynamic target is determined based on the interval and the offset; The velocity and heading angle are optimized according to the second preset loss function so that the dynamic target can be detected by the optimized second detection head and the motion state of the dynamic target can be predicted.
5. The method of claim 2, wherein, The third objective is to identify a safe driving area; the third detection head includes a second encoder and a first decoder. The method for training and optimizing the third detection head based on the second feature map includes: The target segmentation map is output based on the second encoder, the first decoder, and the second feature map. Based on the target segmentation map and the preset line-surface joint learning algorithm, the binary segmentation map of the safe driving area and the edge map of the obstacle are output. The parameters of the binary segmentation map and the edge map are optimized according to the third preset loss function so that the safe driving area can be identified by the optimized third detection head.
6. The method of claim 1, wherein, After acquiring the target image data, the process also includes: The target image data is annotated and projected using a preset annotation tool and a preset inverse perspective transformation algorithm to obtain the bird's-eye view planar coordinate representation of the target image data in the vehicle coordinate system; A target format annotation file is generated based on the target image data represented by the planar coordinates of the bird's-eye view.
7. The method of claim 1, wherein, The vehicle also includes a point cloud acquisition module; The training method also includes: In the initial stage of training the target perception model, point cloud data collected by the point cloud acquisition module is obtained, and the target perception model is trained and optimized based on the point cloud data. 8.A training system of a perception model, characterized in that, An automatic parking system for vehicles, wherein the vehicle includes an image acquisition module; the training system includes: The first acquisition module is used to acquire target image data; wherein, the target image data is two-dimensional image data of the vehicle in the target parking scene acquired by the image acquisition module; The second acquisition module is used to acquire the target perception model; the target perception model is based on a multi-task bird's-eye view perception model. The training module is used to train and optimize the target perception model based on the target image data to obtain a trained target perception model; wherein the trained target perception model is used to perceive multiple target tasks.
9. An electronic device, comprising: It includes a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements the method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method as described in any one of claims 1-7.