A method, system and device for intelligent driving training enhancement based on multi-viewing
Through the multi-camera surround view system and SLAM technology, obstacles and marking lines can be identified in real time, a bird's-eye view of the virtual site can be generated, and real-time feedback and intervention can be provided, which solves the safety risks and lack of authenticity of the driving training system and improves the training effect.
Patent Information
- Application Number
- CN202310295573.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-22
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-03-22
AI Technical Summary
The existing driving training system has safety risks, lacks authenticity, and provides a poor user experience. Differences in instructor levels and trainees' learning abilities lead to widely varying learning outcomes, and the electronic training system does not provide a realistic experience.
A multi-camera surround view system combined with deep learning and SLAM technology is used to collect panoramic images in real time, identify obstacles and marking lines, generate a bird's-eye view of the virtual site, provide real-time feedback and intervention, obtain vehicle status through inertial sensors, generate the optimal motion trajectory and visualize it.
It improves the safety and authenticity of driving training, reduces the risk of accidents, enhances students' dynamic perception and error correction capabilities, provides intuitive operating guidance, and improves training efficiency.
Smart Images

Figure CN116311131B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing technology, and specifically relates to an intelligent driving training enhancement method and system based on multi-view surround vision, and corresponding devices. Background Art
[0002] Motor vehicle drivers must possess a valid driver's license to operate their vehicles on the road. Drivers are only allowed to apply for and obtain a driver's license after completing relevant traffic knowledge and driving skills training, and passing an assessment conducted by the relevant authorities. For most drivers, receiving driving skills training at a driving school is almost a prerequisite for obtaining a driver's license.
[0003] Currently, driving training schools generally employ a one-on-one or one-on-multiple instructional model between instructors and students. After students have acquired a certain level of vehicle knowledge, instructors typically let them practice driving and provide guidance. However, due to differences in instructors' teaching methods and skills, and even greater variations in students' learning abilities, applying the same training methods can not only lead to widely varying learning outcomes but can even result in serious traffic accidents. Therefore, driving schools should not allow students with no prior vehicle operation experience to operate vehicles in open environments or enclosed areas where people are present.
[0004] To overcome this problem, some driving schools often deploy electronic training systems, essentially simulators that simulate vehicle driving maneuvers. These systems consist of a 1:1 simulated cockpit and an electronic screen. As the trainee performs driving maneuvers in the simulated cockpit, the screen changes accordingly, helping them master the operating logic of key vehicle components like the accelerator, clutch, brake, and gear lever. However, the experience offered by these systems is still not realistic enough, significantly differing from the training provided by actual driving. Summary of the Invention
[0005] In order to solve the problems of safety risks in real-car training during driving training, insufficient authenticity of electronic training, and poor user experience; the present invention provides an intelligent driving training enhancement method and system based on multi-view surround vision, as well as corresponding devices.
[0006] The present invention is achieved by adopting the following technical solutions:
[0007] A multi-camera surround vision-based intelligent driving training enhancement method is used to visualize the driving behavior of driving school students to enhance their dynamic perception of the environment and their ability to correct errors during driving training. The intelligent driving training enhancement method includes the following steps:
[0008] S1: Multiple cameras are installed around the driving vehicle to build a multi-view system that can collect image information corresponding to the vehicle's omnidirectional field of view.
[0009] S2: Establish an obstacle recognition network and a landmark recognition network based on deep learning, including the following steps:
[0010] S21: A large number of panoramic images of the real vehicle surrounding environment are collected through a multi-camera surround view system, and the panoramic images are cropped to form a sample dataset.
[0011] S22: Construct an image recognition network based on convolutional neural network, BP neural network or generative adversarial network to achieve target feature recognition.
[0012] S23: Classify and manually label the sample data set to obtain an obstacle training set and a landmark training set respectively.
[0013] S24: Preset training parameters and use the obstacle training set and the landmark training set to train and test the image recognition network respectively.
[0014] S25: retain the network parameters of the two types of network models with the best performance after training, and obtain the required obstacle recognition network and landmark recognition network.
[0015] S3: Determine the type of training task for the trainee and obtain the panoramic image data collected by the multi-camera surround view system in real time. Then, use SLAM technology to create a virtual three-dimensional model of the site based on the panoramic image data, and combine the training task type and the three-dimensional model of the site to generate a bird's-eye view of the site.
[0016] S4: Using the obstacle recognition network and the marker line recognition network to synchronously perform feature recognition on the image data acquired in real time in step S3, and extract the obstacles and marker lines contained therein.
[0017] S5: After image correction, the obstacles and marking lines extracted in step S4 are mapped back to the created virtual three-dimensional model of the site, and the marking lines and obstacles are projected onto the generated bird's-eye view of the site.
[0018] S6: Acquire the data collected by the inertial sensor during vehicle driving, and combine the vehicle's inertial motion data and the image data of the multi-camera surround view system to achieve vehicle positioning; generate the vehicle's real-time trajectory and real-time movement direction through a polynomial algorithm based on the relative changes in the vehicle's position.
[0019] S7: During the vehicle driving process, a reference motion trajectory range of the vehicle that satisfies the collision-free constraint is generated based on the known positions of obstacles and marking lines; a motion trajectory that best matches the current real-time trajectory is designated as the optimal motion trajectory, and a desired motion direction of the optimal motion trajectory is generated.
[0020] S8: During the driving process, the driver is shown changes in the field of view of the bird's-eye view of the site through any visualization method, and the real-time trajectory and real-time movement direction of the vehicle, as well as the corresponding movement trajectory range, optimal movement trajectory and expected movement direction, are dynamically displayed in the field bird's-eye view; and real-time feedback and / or active intervention are provided to the trainee on incorrect driving behaviors.
[0021] As a further improvement of the present invention, in step S1, the multi-camera surround view system includes an image acquisition unit and a data processing unit. The image acquisition unit includes three monocular cameras with a 120° field of view and one monocular camera with a 90° field of view. The first three monocular cameras are positioned in front, on the left, and on the right side of the vehicle, respectively; the last monocular camera is positioned at the rear of the vehicle. The image processing unit is configured to first perform distortion correction on the image data captured by the four monocular cameras based on pre-calibrated camera parameters; then, using image registration and image fusion techniques, the corrected images are stitched together into a seamless panoramic image.
[0022] As a further improvement of the present invention, in step S22, during the image recognition network construction process, the backbone, connections, and branches of the network architecture are first determined based on the model's functionality, performance, and accuracy requirements. The network architecture is then built using the PyTorch framework. Finally, the network model quantifies the image content based on the image features of the input sample image. The image features of the sample image include shape, texture, and color.
[0023] As a further improvement of the present invention, in step S3, the training task types of the trainees are divided into two categories: site tasks and road tasks; the site tasks are further divided into five sub-items: driving on curves, turning at right angles, side parking, parking at a fixed point on a slope, and reversing into a garage.
[0024] The established bird's-eye view of the site includes a solid modeling part generated based on the panoramic image and a virtual modeling part expanded based on the complete site or road; at the same time, the solid modeling part in the bird's-eye view of the site synchronously marks the identified obstacles and marking lines.
[0025] As a further improvement of the present invention, in step S6, the inertial motion data of the vehicle includes the vehicle speed and front wheel angle obtained by the vehicle control system, and the vehicle acceleration collected by the inertial sensor.
[0026] As a further improvement of the present invention, in step S8, when there is a risk of collision between the obtained real-time driving trajectory of the vehicle and an obstacle, an early warning is issued to the driver; when the tires and the marking line overlap in the real-time driving trajectory of the vehicle, the driver's wrong behavior is recorded.
[0027] As a further improvement of the present invention, the criteria for determining whether the driver has engaged in erroneous driving behavior include:
[0028] a. During reverse parking and parallel parking, check whether the vehicle's tires are on the line at the parking position, whether the vehicle body has swept the line during movement, and whether the vehicle body is within the specified area after parking.
[0029] b. During the slope parking project, the driver's operating procedures are correct, the distance between the vehicle tires and the specified line exceeds the preset limit, and the vehicle slides down the slope.
[0030] c. During right-angle turns, check whether the vehicle's tires are on the line and whether the vehicle body is moving along the line.
[0031] d. During the curve driving project, check whether the vehicle tires are on the line and whether the vehicle body is on the line.
[0032] As a further improvement of the present invention, in step S8, the device for visually presenting the bird's-eye view of the site to the driver uses the vehicle's central control screen, or a HUD display installed on the vehicle's front windshield, or a smart wearable device with display function worn by the driver.
[0033] The present invention also includes an intelligent driving training enhancement system based on multi-camera surround vision, which uses the aforementioned intelligent driving training enhancement method based on multi-camera surround vision to visualize the driver's driving behavior and provide driving behavior guidance to the driver. The intelligent driving training enhancement system includes: a multi-camera assembly, an inertial sensor, a data processing module, a display, and a storage module.
[0034] The multi-camera assembly includes three monocular cameras with a 120° field of view and one monocular camera with a 90° field of view. The first three monocular cameras are positioned at the front, left, and right sides of the vehicle; the last monocular camera is positioned at the rear. Inertial sensors measure the vehicle's linear acceleration during operation.
[0035] The data processing module is electrically connected to the multi-camera assembly and the inertial sensor, and the data processing module is also communicatively connected to the vehicle control system.
[0036] The data processing module runs a panoramic image synthesis unit, an obstacle recognition network, a landmark recognition network, a SLAM modeling unit, a site bird's-eye view map generation unit, a motion trajectory generation unit, and a path planning unit. The panoramic image synthesis unit first performs distortion correction on the image data captured by the four monocular cameras based on pre-calibrated camera parameters. The corrected images are then stitched together into a seamless panoramic image using image registration and image fusion techniques. The obstacle recognition network identifies and extracts obstacles from the input panoramic image. The landmark recognition network identifies and extracts ground landmarks from the input panoramic image. The SLAM modeling unit creates a virtual three-dimensional site model based on the input global image. The site bird's-eye view map generation unit projects the identified landmarks and obstacles onto the generated site bird's-eye view map. The motion trajectory generation unit acquires data collected by the vehicle's inertial sensors during driving, combining the vehicle's inertial motion data with image data from the multi-camera surround view system to achieve vehicle positioning. Based on the relative changes in vehicle position, a polynomial algorithm is then used to generate the vehicle's real-time trajectory and direction of motion. The path planning unit is used to generate a reference motion trajectory range of the vehicle that meets the collision-free constraint based on the known positions of obstacles and marking lines; and designate the motion trajectory that best matches the current real-time trajectory as the optimal motion trajectory, and at the same time generate the expected motion direction of the optimal motion trajectory.
[0037] The display is electrically connected to the data processing module, and the display is used to display in real time the bird's-eye view of the site output by the data processing module, including the real-time trajectory and real-time movement direction of the vehicle, as well as the corresponding movement trajectory range, optimal movement trajectory and expected movement direction.
[0038] The storage module is used to save the dynamic images displayed on the display during the driving process, and record the corresponding collision warnings and marking information of the student's incorrect driving behavior according to the timeline.
[0039] The present invention also includes a multi-camera surround-view intelligent driving training enhancement device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor. The multi-camera surround-view intelligent driving training enhancement device is designed according to the architecture of the data processing module in the aforementioned multi-camera surround-view intelligent driving training enhancement system. When the multi-camera surround-view intelligent driving training enhancement device is installed on a vehicle equipped with a multi-camera assembly and an inertial sensor assembly, the processor executes the computer program, implementing the steps of the aforementioned multi-camera surround-view intelligent driving training enhancement method, thereby utilizing the vehicle's existing or installed display module to provide operational guidance services to students receiving driving training.
[0040] The technical solution provided by the present invention has the following beneficial effects:
[0041] The present invention uses a multi-camera surround view system to collect panoramic images of the vehicle's circumference, and uses SLAM technology to model the vehicle's surrounding environment based on the collected panoramic images. The present invention also generates a bird's-eye view of the training site where the vehicle is located, combining typical spaces of different training scenarios. Obstacles and marking lines identified by image recognition technology are also marked in the bird's-eye view of the site. In addition, the present invention combines the vehicle's inertial state data with the spatial position relationship of the vehicle determined by panoramic images, and then solves the vehicle's precise position and motion state in the training site. Based on the continuous change of the vehicle's spatial position, the present invention also realizes real-time tracking and posture prediction of the vehicle; and ultimately realizes the verification of the vehicle's driving route and the generation of the best path to guide the driver's operation.
[0042] The bird's-eye view of the site and the various vehicle status data generated by this invention can ultimately be presented to the driver using conventional display technology or augmented reality technology, allowing the driver to dynamically perceive the vehicle's accurate spatial position while driving, thereby improving the effectiveness of driving training. Furthermore, the invention also uses the acquired vehicle spatial information to provide early warnings and proactive intervention for dangerous vehicle conditions during driving, thereby reducing the risk of serious collisions during training.
[0043] This invention uses SLAM technology to simultaneously acquire site information and identify the vehicle's position and track its trajectory. Compared to conventional driving training systems on the market that use GPS positioning, this solution no longer relies on satellite or radar signals, resulting in greater stability in venues such as on-site or indoor locations. This solution is both innovative and more practical. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0045] Figure 1 This is a flowchart of the steps of an intelligent driving training enhancement method based on multi-view surround vision provided in Example 1 of the present invention.
[0046] Figure 2 This is a case effect diagram of image correction using the chessboard calibration method in Example 1 of the present invention.
[0047] Figure 3 This is a motion trajectory planning diagram when the solution provided in Example 1 of the present invention is applied in a parallel parking training program.
[0048] Figure 4 This is a system module block diagram of an intelligent driving training enhancement system based on multi-viewing provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0050] Example 1
[0051] The present embodiment provides an intelligent driving training enhancement method based on multi-eye surround vision, which is used to visualize the driving behavior of driving school students, so as to enhance the students' dynamic perception of the environment and behavior error correction capabilities during driving training. Specifically, the technical solution provided by this embodiment mainly includes the following key technical points. 1. Real-time collection of image data around the vehicle through a multi-eye surround vision system. 2. Use the collected image data to build a three-dimensional virtual model of the vehicle's surroundings and identify the obstacles contained therein. 3. Generate a bird's-eye view of the site around the vehicle based on the three-dimensional virtual model. 4. Track and display the vehicle's real-time position, trajectory and other status data in the generated bird's-eye view of the site, and provide relevant feedback such as operation guidance and safety warnings based on the driver's driving behavior.
[0052] Specifically, if Figure 1 As shown, the intelligent driving training enhancement method provided in this embodiment includes the following steps:
[0053] S1: Install multiple cameras around the driving vehicle to build a multi-eye surround view system that can collect image information corresponding to the vehicle's omnidirectional field of view. The multi-eye surround view system includes an image acquisition unit and a data processing unit. The angle that a single camera can capture is limited. In order to build a 360° surround view system, it is necessary to capture images through multiple groups of cameras and perform image stitching to achieve the desired effect. The image acquisition unit in this embodiment includes three monocular cameras with a field of view of 120° and one monocular camera with a field of view of 90°. The first three monocular cameras are respectively arranged in front, left and right of the vehicle; the last monocular camera is arranged at the rear of the vehicle.
[0054] The image processing unit first performs distortion correction on the image data captured by the four monocular cameras based on pre-calibrated camera parameters. The corrected images are then stitched together into a seamless panoramic image using image registration and image fusion techniques. For this image correction task, this embodiment, based on ROS in Ubuntu (version noetic), utilizes a checkerboard calibration method to obtain camera intrinsic and extrinsic parameters. The calibration process is continuously progressed as much as possible, and finally, distortion correction is performed on the images through the calibration process to obtain the target video image. This method improves calibration accuracy and also enhances the system's distortion correction capabilities. Figure 2This is a typical case of image correction achieved using the method of this embodiment.
[0055] Image stitching technology is a technology that stitches together several images with overlapping parts (corresponding to the images captured by multiple different monocular cameras in this embodiment) into a seamless panoramic image or high-resolution image. Among them, image registration and image fusion are two key technologies for image stitching. In this embodiment, image registration technology can identify the overlapping parts in images from different sources, and then select this part as the benchmark for the overlapping parts in the stitching process of different images. In VS (2019 version) + opencv, there are already built-in functions and a built-in stitching algorithm stitch. These codes can be used to realize feature point selection and save the optimal feature point object; feature point matching, perspective transformation and image fusion, and finally the optimized image is displayed as the final result to obtain a 360-degree panoramic view image around the vehicle body. The spliced 360-degree panoramic view image obtained in this embodiment can be used as a real-time input for subsequent target recognition and spatial modeling.
[0056] S2: This embodiment requires real-time analysis of the position information of obstacles and marking lines around the vehicle during the vehicle position tracking and trajectory optimization process. To this end, this embodiment specially constructs an image recognition network that can be used to identify various different targets. These image recognition networks are mainly used to achieve the purpose of identifying and locating various obstacles and marking lines contained in the collected 360-degree panoramic images. In the process of building the image recognition network, the backbone, connections and branches of the network architecture are first determined based on the model function, performance and accuracy requirements, and then the network architecture is built using the pytorch framework. Finally, the network model quantifies the image content based on the image features of the input sample image. Among them, the image features of the sample image include shape, texture, and color.
[0057] For the specific application scenario of this embodiment, an obstacle recognition network and a marker line recognition network based on deep learning are established, specifically including the following steps:
[0058] S21: A large number of panoramic images of the real vehicle surrounding environment are collected through a multi-camera surround view system, and the panoramic images are cropped to form a sample dataset.
[0059] S22: Construct an image recognition network based on convolutional neural network, BP neural network or generative adversarial network to achieve target feature recognition.
[0060] S23: Classify and manually label the sample data set to obtain an obstacle training set and a landmark training set respectively.
[0061] In this example, we used the video calibration software Labelme to calibrate the processed video image data. Lane lines were distinguished using different naming schemes, such as lane1, lane2, and so on. Next, different attributes appearing in the video, such as solid lines and dashed lines, needed to be categorized and labeled, distinguished by different names. Calibration was completed by improving the output frame by frame.
[0062] S24: Preset training parameters and use the obstacle training set and the landmark training set to train and test the image recognition network respectively.
[0063] During the training of the network model in this embodiment, a dataset loader is constructed based on the dataset type used to preprocess and load the raw data. During model usage, PyTorch automatically calls a function to perform parameter forward propagation. The inherited data function reads the specified data from the passed dataset path and uses the function to centralize, standardize, and normalize the dataset. The initial dataset is then split to facilitate subsequent evaluation.
[0064] Next, determine the optimizer used for the network model. Use a storage system with sufficient storage capacity and a GPU capable of handling deep learning data modeling computations. Use optimizers with different characteristics to meet different requirements (such as SGD and Adam). Update the model parameter file and load the original dataset into the GPU. Iterate and train the data in the dataset object as needed, passing it into the specified statement. Then, based on the PyTorch framework, evaluate and adjust the accuracy of the test or validation set. Construct a network loss function and introduce a loss function to minimize the discrepancy between predicted and actual results.
[0065] S25: retain the network parameters of the two types of network models with the best performance after training, and obtain the required obstacle recognition network and landmark recognition network.
[0066] S3: Determine the trainee's training task type and acquire panoramic image data from the multi-camera surround view system in real time. SLAM technology is then used to create a virtual 3D model of the site based on the panoramic image data. Combined with the training task type and the 3D site model, a bird's-eye view of the site is generated. This bird's-eye view includes a solid modeling component generated from the panoramic image and a virtual modeling component expanded from the complete site or road. The solid modeling component of the bird's-eye view also marks identified obstacles and marking lines.
[0067] The solution provided in this embodiment primarily addresses common driving school inspection scenarios, such as cornering, right-angle turns, parallel parking, parking on ramps, and reversing into a garage. It integrates the display and correction of vehicle driving trajectories and provides targeted optimization. The training tasks for students are divided into two categories: field tasks and road tasks. Field tasks correspond to the various items in the Subject II exam, while road tasks target the various items in Subject III. Considering that students in Subject III already possess basic driving skills, while students in Subject II generally have a lower level of driving proficiency, posing a greater risk in the driving learning process, this solution focuses more on optimizing and adapting field tasks.
[0068] The bird's-eye view of the site provided in this example is equivalent to providing a bird's-eye view of a camera that can follow the movement of the vehicle in real time and shoot from the top of the vehicle. This perspective has the most intuitive and accurate guiding effect on the vehicle. Taking into account that the panoramic image obtained by the multi-eye surround view system cannot completely cover the entire site during the training process, this embodiment combines real-scene modeling and virtual modeling. For the close-up part of the vehicle, the panoramic image is used for modeling, and the identified obstacles and marking lines are marked. For the part that the global image does not cover, virtual modeling is used to "complete" it; thus presenting the driver with a complete bird's-eye view of the site. The solution provided in this embodiment is mainly designed for site tasks, and the site ranges of different training items in the site tasks are mostly fixed, which is easy to achieve in the virtual modeling process.
[0069] S4: Using the obstacle recognition network and the marker line recognition network to synchronously perform feature recognition on the image data acquired in real time in step S3, and extract the obstacles and marker lines contained therein.
[0070] S5: After image correction, the obstacles and marking lines extracted in step S4 are mapped back to the created virtual three-dimensional model of the site, and the marking lines and obstacles are projected onto the generated bird's-eye view of the site.
[0071] In this embodiment, the essence of marking obstacles and marking lines in the bird's-eye view of the site is to "reproject" the information of the target objects identified in the panoramic image onto the ground plane. However, the image obtained by the camera's oblique view will generally be distorted. Therefore, in the process of projecting marking lines and obstacles, this embodiment first needs to obtain the projection transformation relationship H between the plane corresponding to the camera's field of view and the bottom surface. The specific method is: by placing the calibration plate image on the ground plane, the coordinates of the four vertices on the checkerboard image of the ground plane are obtained: (0, 0), (widht-1, 0), (0, height-1), (wdith-1, height-1); at the same time, the corner points of the captured image plane are extracted, and the coordinate values of the corner points corresponding to the four points on the ground plane in the image space are obtained; finally, through the correspondence between the four coordinate points, the projection transformation relationship H between the ground plane and the image plane corresponding to the camera's field of view is obtained based on the function.
[0072] After determining the projection transformation relationship H, the image is inversely mapped to the ground plane space using a function. In VS (2019) + OpenCV, there is an embedded function that can be called by code to complete the top view conversion operation of the image of the identified target area.
[0073] S6: Acquire data collected by inertial sensors during vehicle driving, combine the vehicle's inertial motion data with image data from the multi-camera surround view system to achieve vehicle positioning; and generate the vehicle's real-time trajectory and direction of motion based on the relative changes in the vehicle's position using a polynomial algorithm. The vehicle's inertial motion data includes vehicle speed and front wheel angle obtained by the vehicle control system, as well as vehicle acceleration collected by the inertial sensors.
[0074] Specifically, the solution provided in this embodiment utilizes a vehicle-mounted multi-camera surround view system to capture panoramic vision around the vehicle. Combined with SLAM technology, this system can create a three-dimensional model of the training site. It then collects and pre-processes vehicle inertial sensor data, followed by camera image data. This combined data allows for positioning, determining the vehicle's real-time relative position.
[0075] S7: During the vehicle driving process, a reference motion trajectory range of the vehicle that satisfies the collision-free constraint is generated based on the known positions of obstacles and marking lines; a motion trajectory that best matches the current real-time trajectory is designated as the optimal motion trajectory, and a desired motion direction of the optimal motion trajectory is generated.
[0076] S8: During driving, the system uses any visualization method to display changes in the bird's-eye view of the site to the driver. The system dynamically displays the vehicle's real-time trajectory and direction of movement, along with the corresponding trajectory range, optimal trajectory, and desired direction of movement. The system also provides real-time feedback and / or proactive intervention for incorrect driving behavior. If there is a risk of collision between the vehicle's real-time trajectory and an obstacle, the system issues a preemptive warning to the driver. If the tires of the vehicle overlap a marking line within the real-time trajectory, the system records the driver's incorrect behavior.
[0077] The following is a brief introduction to the application process of this embodiment in the parallel parking project with reference to the accompanying drawings to highlight the advantages of this embodiment:
[0078] During training, the system provides an optimal parallel parking trajectory for the current position. The screen also displays real-time vehicle position information through a bird's-eye view of image stitching. Comparing this information with the optimal trajectory allows trainees to make immediate corrections. Through repeated driving practice, trainees can acquire proficient parallel parking techniques.
[0079] The vehicle trajectory image obtained by parallel parking is as follows Figure 3 As shown, by calibrating the vehicle's current position, obstacles, and parking lot markings, the boundary constraints during driving can be determined. By analyzing these constraints, safety boundary constraint data is obtained, which is then incorporated into a polynomial calculation to obtain a simulated image of the vehicle's optimal motion trajectory. Combined with the bird's-eye view on the display, the optimal trajectory is blurred and compared with the real-time position to help the driver determine a correct parking practice route.
[0080] In addition to parallel parking, the present invention's application process is similar for training exercises such as cornering, right-angle turns, parallel parking, parking on ramps, and reversing into parking spaces. By analyzing the constraints of various obstacles (road markings and boundaries), the correct driving trajectory can be determined. This allows for a correct driving experience in the early stages of practice.
[0081] Compared to traditional instructors' verbal instruction, this embodiment provides a more intuitive on-screen display of driving trajectories and correction methods during the early stages of training. By practicing the correct trajectories in the early stages, students can develop a sense of autonomous driving more quickly, improving the efficiency of driving training.
[0082] During application, this embodiment can accurately determine whether a driver has engaged in erroneous driving behavior based on preset conditions. For example, in a routine field mission, the criteria for determining whether a driver has engaged in erroneous driving behavior (driving behavior that violates regulations but does not cause personal injury or property damage) include:
[0083] a. During reverse parking and parallel parking, check whether the vehicle's tires are on the line at the parking position, whether the vehicle body has swept the line during movement, and whether the vehicle body is within the specified area after parking.
[0084] b. During the slope parking project, the driver's operating procedures are correct, the distance between the vehicle tires and the specified line exceeds the preset limit, and the vehicle slides down the slope.
[0085] c. During right-angle turns, check whether the vehicle's tires are on the line and whether the vehicle body is moving along the line.
[0086] d. During the curve driving project, check whether the vehicle tires are on the line and whether the vehicle body is on the line.
[0087] Furthermore, the solution provided in this embodiment can also send instructions to the vehicle safety control system when a student engages in behavior that poses a risk of personal injury or property damage, allowing the vehicle safety system to provide a warning or proactively intervene in the student's driving behavior. For example, if a student mistakenly steps on the accelerator, causing the vehicle to collide with an obstacle while moving forward or backward, the solution of the present invention can proactively detect the approaching obstacle in the bird's-eye view of the scene and issue a braking control instruction to the vehicle safety control system before the instructor in the passenger seat does.
[0088] In this embodiment, the driver can receive a real-time bird's-eye view of the vehicle's dynamic changes in the venue presented by visualization means during driving. The bird's-eye view of the venue will also display vehicle status data such as real-time trajectory and optimal motion trajectory. These images can be displayed on the vehicle's existing central control screen. Currently, many high-performance smart cars have replaced the instrument panel with an electronic screen. This embodiment can also display the relevant images on the central control screen. In addition, in order to prevent the rear display screen from interfering with the driver's visual concentration and operating behavior, the vehicle can also be equipped with a special HUD head-up display, which can present the required various types of driving assistance information on the front windshield of the vehicle through the HUD display. Alternatively, the required driving assistance information can be displayed through various smart wearable devices with display functions worn by the driver (such as AR glasses, etc.) to guide the driver's vehicle driving behavior.
[0089] It should be emphasized that: this embodiment adopts a multi-view surround view system + SLAM + image-based target recognition technology to achieve site information acquisition, while also achieving vehicle posture recognition and trajectory tracking; and on this basis, it realizes assistance and guidance of the driver's driving behavior. In conventional driving training systems on the market, obstacle detection technology that combines GPS technology and lidar can also solve the same problem. However, compared with the GPS+lidar solution, the solution of this embodiment no longer relies on satellite signals or lidar information, and therefore has higher stability and reliability in venues or indoors, and the solution of the present invention has better recognition effect on dynamic targets than lidar. Therefore, the solution of this embodiment is more innovative and more practical.
[0090] Example 2
[0091] This embodiment provides an intelligent driving training enhancement system based on multi-view surround vision. The system is a complete set of software and hardware solutions. The solution mainly adopts the intelligent driving training enhancement method based on multi-view surround vision as in Example 1 to visualize the driver's driving behavior and provide driving behavior guidance to the driver.
[0092] like Figure 4 As shown, the intelligent driving training enhancement system provided in this embodiment includes: a multi-camera assembly, an inertial sensor, a data processing module, a display, and a storage module.
[0093] The multi-camera assembly includes three monocular cameras with a 120° field of view and one monocular camera with a 90° field of view. The first three monocular cameras are positioned at the front, left, and right sides of the vehicle; the last monocular camera is positioned at the rear. Inertial sensors measure the vehicle's linear acceleration during operation.
[0094] The data processing module is electrically connected to the multi-camera assembly and the inertial sensor, and the data processing module is also communicatively connected to the vehicle control system.
[0095] The data processing module runs a panoramic image synthesis unit, an obstacle recognition network, a landmark recognition network, a SLAM modeling unit, a site bird's-eye view map generation unit, a motion trajectory generation unit, and a path planning unit. The panoramic image synthesis unit first performs distortion correction on the image data captured by the four monocular cameras based on pre-calibrated camera parameters. The corrected images are then stitched together into a seamless panoramic image using image registration and image fusion techniques. The obstacle recognition network identifies and extracts obstacles from the input panoramic image. The landmark recognition network identifies and extracts ground landmarks from the input panoramic image. The SLAM modeling unit creates a virtual three-dimensional site model based on the input global image. The site bird's-eye view map generation unit projects the identified landmarks and obstacles onto the generated site bird's-eye view map. The motion trajectory generation unit acquires data collected by the vehicle's inertial sensors during driving, combining the vehicle's inertial motion data with image data from the multi-camera surround view system to achieve vehicle positioning. Based on the relative changes in vehicle position, a polynomial algorithm is then used to generate the vehicle's real-time trajectory and direction of motion. The path planning unit is used to generate a reference motion trajectory range of the vehicle that meets the collision-free constraint based on the known positions of obstacles and marking lines; and designate the motion trajectory that best matches the current real-time trajectory as the optimal motion trajectory, and at the same time generate the expected motion direction of the optimal motion trajectory.
[0096] The display is electrically connected to the data processing module, and the display is used to display in real time the bird's-eye view of the site output by the data processing module, including the real-time trajectory and real-time movement direction of the vehicle, as well as the corresponding movement trajectory range, optimal movement trajectory and expected movement direction.
[0097] The storage module is used to save the dynamic images displayed on the display during the driving process, and record the corresponding collision warnings and marking information of the student's incorrect driving behavior according to the timeline.
[0098] Example 3
[0099] This embodiment provides an intelligent driving training enhancement device based on multi-eye surround vision, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. The intelligent driving training enhancement device based on multi-eye surround vision is designed according to the architecture of the data processing module in the intelligent driving training enhancement system based on multi-eye surround vision in Example 2. When the intelligent driving training enhancement device based on multi-eye surround vision is installed on a vehicle equipped with a multi-eye camera assembly and an inertial sensor assembly. When the processor executes the computer program, the steps of the intelligent driving training enhancement method based on multi-eye surround vision in Example 1 are implemented, and then the original or additional display module of the vehicle is used to provide operation guidance services for students receiving driving training.
[0100] The intelligent driving training enhancement device in this embodiment is actually a computer device. The computer device can be a smart terminal capable of executing programs, a tablet computer, a laptop computer, a desktop computer, a rack server, a blade server, a tower server, or a cabinet server (including a standalone server or a server cluster consisting of multiple servers). The computer device in this embodiment includes at least, but is not limited to, a memory and a processor that can be interconnected via a system bus.
[0101] In this embodiment, the memory (i.e., readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic storage, magnetic disk, optical disk, etc. In some embodiments, the memory may be an internal storage unit of a computer device, such as the hard disk or internal memory of the computer device. In other embodiments, the memory may also be an external storage device of the computer device, such as a plug-in hard disk equipped with the computer device, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the memory may also include both the internal storage unit of the computer device and its external storage device. In this embodiment, the memory is generally used to store the operating system and various application software installed on the computer device. In addition, the memory may also be used to temporarily store various types of data that have been output or are about to be output.
[0102] In some embodiments, the processor may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor is generally used to control the overall operation of a computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data to implement the steps of the intelligent driving training enhancement method based on multi-view in the aforementioned embodiment 1, thereby generating the required bird's-eye view of the site containing a variety of information, and using the vehicle's original or added display module to provide operation guidance services for students receiving driving training.
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. An intelligent driving training enhancement method based on multi-viewing, characterized in that: It is used to visualize the driving behavior of driving school students to enhance the students' dynamic perception of the environment and behavior error correction capabilities during driving training. The intelligent driving training enhancement method includes the following steps: S1: Multiple cameras are installed around the driving vehicle to build a multi-view system that can collect image information corresponding to the vehicle's omnidirectional field of view; S2: Establish an obstacle recognition network and a landmark recognition network based on deep learning, including the following steps: S21: A large number of panoramic images of the real vehicle surrounding environment are collected through a multi-camera surround view system, and the panoramic images are cropped to form a sample dataset; S22: Construct an image recognition network based on convolutional neural network, BP neural network or generative adversarial network to achieve target feature recognition; S23: Classify and manually label the sample data set to obtain an obstacle training set and a landmark training set respectively; S24: Preset training parameters and use the obstacle training set and the landmark training set to train and test the image recognition network respectively; S25: retain the network parameters of the two network models with the best performance after training, and obtain the required obstacle recognition network and landmark recognition network; S3: Determine the type of training task for the trainee and obtain the panoramic image data collected by the multi-camera surround view system in real time. Then, use SLAM technology to create a virtual 3D model of the site based on the panoramic image data. Combined with the training task type and the 3D model of the site, a bird's-eye view of the site is generated. S4: using the obstacle recognition network and the marker line recognition network to synchronously perform feature recognition on the image data acquired in real time in step S3, and extracting obstacles and marker lines contained therein; S5: Perform image correction on the obstacles and marking lines extracted in step S4 and then map them back to the created virtual three-dimensional model of the site, and project the marking lines and obstacles onto the generated bird's-eye view of the site; S6: Acquire data collected by inertial sensors during vehicle driving, combine the vehicle's inertial motion data with image data from the multi-camera surround view system to achieve vehicle positioning; generate the vehicle's real-time trajectory and real-time movement direction based on the relative changes in the vehicle's position using a polynomial algorithm; S7: During the vehicle driving process, a reference motion trajectory range of the vehicle that satisfies the collision-free constraint is generated based on the known positions of obstacles and marking lines; a motion trajectory that best matches the current real-time trajectory is designated as the optimal motion trajectory, and a desired motion direction of the optimal motion trajectory is generated; S8: During the driving process, the changes in the field of view of the bird's-eye view of the site are displayed to the driver through any visualization method, and the real-time trajectory and real-time movement direction of the vehicle, as well as the corresponding movement trajectory range, optimal movement trajectory and expected movement direction are dynamically displayed in the bird's-eye view of the site; and real-time feedback and / or active intervention are provided to the trainee's erroneous driving behavior.
2. The intelligent driving training enhancement method based on multi-viewing according to claim 1, characterized in that: In step S1, the multi-camera surround view system includes an image acquisition unit and a data processing unit; the image acquisition unit includes three monocular cameras with a field of view of 120° and one monocular camera with a field of view of 90°, the first three monocular cameras being respectively arranged in front, on the left, and on the right side of the vehicle; the last monocular camera being arranged at the rear of the vehicle; the image processing unit is used to first perform distortion correction on the image data collected by the four monocular cameras according to pre-calibrated camera parameters; The rectified images are then stitched into a seamless panoramic image based on image registration and image fusion techniques.
3. The intelligent driving training enhancement method based on multi-viewing according to claim 1, characterized in that: In step S22, during the construction of the image recognition network, the backbone, connections, and branches of the network architecture are first determined based on the model function, performance, and accuracy requirements, and then the network architecture is built using the pytorch framework. Finally, the network model quantifies the image content based on the image features of the input sample image; wherein the image features of the sample image include shape, texture, and color.
4. The intelligent driving training enhancement method based on multi-viewing according to claim 1, characterized in that: In step S3, the training task types for the trainees are divided into two categories: field tasks and road tasks. Field tasks are further divided into five sub-items: driving on a curve, turning at right angles, parallel parking, parking at a fixed point on a slope, and reversing into a parking space. The established bird's-eye view of the site includes a solid modeling part generated based on the panoramic image and a virtual modeling part expanded based on the complete site or road; at the same time, the solid modeling part in the bird's-eye view of the site synchronously marks the identified obstacles and marking lines.
5. The intelligent driving training enhancement method based on multi-viewing according to claim 1, characterized in that: In step S6, the inertial motion data of the vehicle includes the vehicle speed and front wheel angle obtained by the vehicle control system, and the vehicle acceleration collected by the inertial sensor.
6. The intelligent driving training enhancement method based on multi-viewing according to claim 1, characterized in that: In step S8, when there is a risk of collision between the obtained real-time driving trajectory of the vehicle and an obstacle, an early warning is issued to the driver; when the tires and the marking line overlap in the real-time driving trajectory of the vehicle, the driver's wrong behavior is recorded.
7. The intelligent driving training enhancement method based on multi-viewing according to claim 1, characterized in that: The criteria for determining whether a driver has engaged in erroneous driving behavior include: a. During reverse parking and parallel parking, check whether the vehicle's tires are on the line at the parking position, whether the vehicle body has crossed the line during movement, and whether the vehicle body is within the specified area after parking; b. During the ramp parking test, the driver's operation procedures are correct, the distance between the vehicle tires and the specified line exceeds the preset limit, and the vehicle rolls down the slope; c. During right-angle turns, check whether the vehicle's tires are on the line and whether the vehicle body is shunting the line during movement; d. During the curve driving project, check whether the vehicle tires are on the line and whether the vehicle body is on the line.
8. The intelligent driving training enhancement method based on multi-viewing according to claim 1, characterized in that: In step S8, the device for visually presenting the bird's-eye view of the site to the driver is a vehicle central control screen, or a HUD display installed on the front windshield of the vehicle, or a smart wearable device with display function worn by the driver.
9. An intelligent driving training enhancement system based on multi-viewing, characterized by: It adopts the intelligent driving training enhancement method based on multi-view according to any one of claims 1 to 8 to visualize the driver's driving behavior and provide driving behavior guidance to the driver; The intelligent driving training enhancement system includes: A multi-camera assembly, comprising three monocular cameras with a 120° field of view and one monocular camera with a 90° field of view. The first three monocular cameras are placed in front, on the left, and on the right side of the vehicle, respectively; the last monocular camera is placed at the rear of the vehicle. Inertial sensors, which are used to measure the linear acceleration of the vehicle during operation; A data processing module is electrically connected to the multi-camera assembly and the inertial sensor, and the data processing module is also communicatively connected to the vehicle control system; a panoramic image synthesis unit, an obstacle recognition network, a marker line recognition network, a SLAM modeling unit, a site bird's-eye view generation unit, a motion trajectory generation unit, and a path planning unit are run in the data processing module; wherein the panoramic image synthesis unit is used to first perform distortion correction on the image data collected by the four monocular cameras according to pre-calibrated camera parameters; and then splice the corrected images into a seamless panoramic image based on image registration and image fusion technology; the obstacle recognition network is used to identify and extract obstacles contained in the input panoramic image; the marker line recognition network is used to identify and extract the site contained in the input panoramic image. surface marker lines; the SLAM modeling unit is used to create a virtual three-dimensional model of the site based on the input global image; the site bird's-eye view generation unit is used to project the identified marker lines and obstacles onto the generated site bird's-eye view; the motion trajectory generation unit is used to obtain data collected by the inertial sensor during vehicle driving, and realize vehicle positioning by combining the inertial motion data of the vehicle and the image data of the multi-camera surround view system; then the real-time trajectory and real-time motion direction of the vehicle are generated by a polynomial algorithm according to the relative change of the vehicle position; the path planning unit is used to generate a reference motion trajectory range of the vehicle that meets the collision-free constraint based on the known positions of obstacles and marker lines; and designate a motion trajectory that best matches the current real-time trajectory as the optimal motion trajectory, and generate the expected motion direction of the optimal motion trajectory; a display electrically connected to the data processing module, the display being configured to display in real time a bird's-eye view of the site output by the data processing module, including the vehicle's real-time trajectory and real-time movement direction, as well as the corresponding movement trajectory range, optimal movement trajectory, and desired movement direction; The storage module is used to save the dynamic images displayed on the display during the driving process and record the corresponding collision warnings and marking information of the student's incorrect driving behavior according to the timeline.
10. An intelligent driving training enhancement device based on multi-viewing, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: The multi-eye surround view intelligent driving training enhancement device is designed according to the architecture of the data processing module in the multi-eye surround view-based intelligent driving training enhancement system according to claim 9. When the multi-eye surround view intelligent driving training enhancement device is installed on a vehicle equipped with a multi-camera assembly and an inertial sensor assembly; when the processor executes the computer program, the steps of the multi-eye surround view-based intelligent driving training enhancement method according to any one of claims 1 to 8 are implemented, and then the original or additional display module of the vehicle is used to provide operation guidance services for students receiving driving training.
Citation Information
Patent Citations
Intelligent driving training system and method based on augment virtual reality man-machine interaction
CN106710360A
Driving test vehicle right-angle turning line pressing detection method and system based on panorama.
CN113538377A