Automatic driving method and system based on machine vision
By combining the vehicle driving speed, flow of people and flow rate in the autonomous driving system, and using the multimodal feature pyramid repair results, the problem of insufficient redundancy and flexibility of semantic segmentation methods in the prior art is solved, and higher accuracy and sensitivity are achieved, and computing power load is reduced.
Patent Information
- Application Number
- CN202510304667.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-14
AI Technical Summary
In the existing autonomous driving technology, the semantic segmentation method is too redundant and cumbersome, resulting in increased computing power load but difficult to improve accuracy. In a highly maneuverable and flexible automotive environment, the flexibility and sensitivity of video stream semantic segmentation are insufficient.
Using an automated driving method based on machine vision, the video stream data is obtained using the vehicle monocular camera, and semantic segmentation is performed based on the vehicle driving speed characteristics, the rate of flow change and the rate of flow change to generate a 3D scene. The method also includes real-time regulating camera shooting parameters and repairing semantic segmentation results using multimodal feature pyramids.
It improves the accuracy and sensitivity of semantic segmentation, enhances the vehicle's perception ability to complex environments, reduces the computing power load of the on-board system, and improves the robustness and adaptability of the system.
Smart Images

Figure CN120182949A_ABST
Abstract
Description
Technical Field
[0001] The present invention provides an autonomous driving method and system based on machine vision, belonging to the technical field of autonomous driving. Background Art
[0002] With the rapid development of artificial intelligence technology, autonomous driving technology has become an important development direction for the future automotive industry. The realization of autonomous driving depends on the comprehensive application of a variety of advanced technologies. Among them, machine vision technology has become one of the core technologies in the field of autonomous driving because it can simulate the functions of the human visual system and achieve accurate perception and understanding of the surrounding environment. In an autonomous driving system, an on-vehicle monocular camera, as the main sensor of machine vision, undertakes the important task of obtaining information about the vehicle's surrounding environment, and determines the drivable area of autonomous driving and obtains the detour path by performing semantic segmentation recognition on the collected video stream data.
[0003] However, in the prior art, in order to improve the accuracy of semantic segmentation during the semantic segmentation process, only the improvement degree of the semantic segmentation method itself is considered, which results in the continuous redundancy and complexity of the semantic segmentation method, increasing the computing power load of the on-vehicle system while being unable to significantly improve the accuracy of semantic recognition. On the other hand, with the continuous development of the performance of new energy vehicles, the mobility and flexibility during vehicle driving are relatively high. In this case, relying solely on the video stream of the vehicle environment for semantic segmentation reduces the flexibility and sensitivity of semantic segmentation. Summary of the Invention
[0004] The present invention provides an autonomous driving method and system based on machine vision to solve the technical problems existing in the above prior art, and the technical solutions adopted are as follows:
[0005] An autonomous driving method based on machine vision, the autonomous driving method based on machine vision includes:
[0006] Obtaining video stream data around the vehicle by using an on-vehicle monocular camera;
[0007] Performing semantic segmentation on the video stream data by using the vehicle driving speed characteristics and the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the on-vehicle monocular camera to obtain a 3D scene after semantic segmentation;
[0008] Performing integrated positioning and detour path generation by using the 3D scene after semantic segmentation and a high-precision map.
[0009] Further, the obtaining video stream data around the vehicle by using an on-vehicle monocular camera includes:
[0010] Real-time regulating the shooting parameters of the on-vehicle monocular camera according to the environmental conditions where the vehicle is located;
[0011] After the shooting parameters of the in-vehicle monocular camera are adjusted, the in-vehicle monocular camera is controlled in real time to perform real-time shooting to obtain the video stream data around the vehicle.
[0012] Furthermore, the video stream data is semantically segmented by using the vehicle driving speed characteristics and the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the in-vehicle monocular camera to obtain a 3D scene after semantic segmentation, including:
[0013] The video stream data is processed by using the vehicle driving speed characteristics and the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the in-vehicle monocular camera to obtain a depth map, a material attribute map, and a 3D scene flow corresponding to the video stream data;
[0014] The 3D scene flow is semantically segmented to obtain an initial semantic segmentation result corresponding to the 3D scene flow;
[0015] Feature extraction and fusion are performed on the continuous RGB video frames, depth map, and material attribute map corresponding to the video stream data to obtain a multi-modal feature pyramid;
[0016] The initial semantic segmentation result is repaired by using the multi-modal feature pyramid and the 3D scene flow to obtain a repaired semantic segmentation result;
[0017] The 3D scene flow is labeled according to the semantic segmentation result to obtain a 3D scene after semantic segmentation.
[0018] Furthermore, the video stream data is processed by using the vehicle driving speed characteristics and the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the in-vehicle monocular camera to obtain a depth map, a material attribute map, and a 3D scene flow, including:
[0019] Frame processing is performed on the video stream data to obtain the continuous RGB video frames corresponding to the video stream data;
[0020] The number of consecutive frames input to the dynamic neural radiance field NeRF model is set according to the maximum and minimum driving speeds of the vehicle within a preset time range in combination with the pedestrian flow change rate and vehicle flow change rate within the current shooting range area of the in-vehicle monocular camera;
[0021] The continuous RGB video frames are divided according to the number of consecutive frames, and the divided continuous RGB video frames are input into the dynamic neural radiance field NeRF model to output a depth map, a material attribute map, and a 3D scene flow corresponding to the video stream data.
[0022] Furthermore, feature extraction and fusion are performed on the continuous RGB video frames, depth map, and material attribute map corresponding to the video stream data to obtain a multi-modal feature pyramid, including:
[0023] Use a convolutional neural network to separately process the semantic feature maps corresponding to consecutive RGB video frames, material property maps, and depth maps;
[0024] Taking the scale of the semantic feature map of the consecutive RGB video as the standard, adjust the semantic feature maps corresponding to the material property map and the depth map to be consistent with the scale of the RGB video semantic feature map, and project the adjusted semantic feature maps corresponding to the material property map and the depth map into the semantic space matching the RGB features for weighted fusion to obtain a multi-modal feature pyramid.
[0025] Furthermore, use the multi-modal feature pyramid and 3D scene flow to repair the initial result of semantic segmentation to obtain the repaired semantic segmentation result, including:
[0026] Use a 3D sparse convolutional kernel to extract features from the 3D voxel grid generated by consecutive RGB video frames to obtain spatio-temporal features;
[0027] Use the current driving speed of the vehicle, the change rate of the pedestrian flow around the vehicle, and the change rate of the vehicle flow, combined with the 3D scene flow and its corresponding density map, to calculate the occlusion area mask of the current frame in the consecutive RGB video frames, and obtain the occlusion area that appears in the current frame;
[0028] Repair the occlusion area to generate the repaired semantic segmentation result.
[0029] Furthermore, the occlusion area mask is 1 or 0. Among them, when the occlusion area mask is 1, it means that the pixel point corresponding to the occlusion area mask is marked as the occlusion area; when the occlusion area mask is 0, it means that the pixel point corresponding to the occlusion area mask is marked as the non-occlusion area.
[0030] Furthermore, the occlusion area mask is obtained by combining the difference between the density value of pixel x in the previous consecutive RGB video frame and the density value of pixel x in the current consecutive RGB video frame, with a dynamically adjusted occlusion threshold and a dynamically adjusted motion residual threshold.
[0031] An autonomous driving system based on machine vision, the autonomous driving system based on machine vision includes:
[0032] A video stream data acquisition module, used to acquire video stream data around the vehicle by using an in-vehicle monocular camera;
[0033] A semantic segmentation module, used to perform semantic segmentation on the video stream data by using the vehicle driving speed characteristics, the change rate of the pedestrian flow, and the change rate of the vehicle flow within the shooting range area of the in-vehicle monocular camera, to obtain a 3D scene after semantic segmentation;
[0034] A path generation module, used to perform integrated positioning and detour path generation by using the 3D scene after semantic segmentation and a high-precision map.
[0035] Advantages of the present invention:
[0036] A machine vision-based autonomous driving method and system provided by the present invention not only rely on the video stream of the vehicle environment during semantic segmentation, but also consider the operating state during vehicle driving and the actual states of road pedestrians and vehicles for semantic segmentation. It can effectively improve the accuracy of semantic segmentation, and to a great extent improve the followability between semantic segmentation and vehicle operation and the automatic adjustability of semantic segmentation. Furthermore, it can maximize the sensitivity and flexibility of semantic segmentation. At the same time, the machine vision-based autonomous driving method and system provided by the present invention can effectively improve the matching between semantic segmentation and vehicle driving state and road state. Thus, when the body is temporarily blocked (such as a pedestrian being briefly blocked by a vehicle and then reappearing), the accuracy and processing efficiency of semantic segmentation can be improved. Also, in the case of rare object materials (such as transparent glass, reflective road surface) or special lighting (such as strong specular reflection) on the road, the accuracy and processing efficiency of semantic segmentation can be improved. At the same time, because the operating state during vehicle driving and the actual states of road pedestrians and vehicles are utilized in the semantic segmentation process, the mobility and sensitivity of the semantic segmentation to process the above special situations can be adjusted according to the operating state during vehicle driving and the actual states of road pedestrians and vehicles. Furthermore, on the premise of ensuring the adaptation of the vehicle operating state to the operation rate of semantic segmentation, while improving the accuracy of intelligent recognition, the computing power of the in-vehicle system can be released to the maximum extent, and the computing power load of the in-vehicle system can be reduced. Description of the drawings
[0037] Figure 1 is a flowchart of the method of the present invention;
[0038] Figure 2 is a system block diagram of the system of the present invention. Detailed implementation manners
[0039] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0040] An embodiment of the present invention proposes a machine vision-based autonomous driving method, as Figure 1 shown, the machine vision-based autonomous driving method includes:
[0041] S1. Obtain the video stream data around the vehicle using an in-vehicle monocular camera; specifically, adjust the shooting parameters of the in-vehicle monocular camera in real time according to the environmental conditions of the vehicle, where the environmental conditions include but are not limited to light intensity and weather conditions; the shooting parameters include but are not limited to resolution, field of view angle, and dynamic range; then, after the shooting parameters of the in-vehicle monocular camera are adjusted, control the in-vehicle monocular camera to perform real-time shooting in real time to obtain the video stream data around the vehicle.
[0042] S2. Perform semantic segmentation on the video stream data using the vehicle driving speed characteristics and the change rates of the pedestrian flow and vehicle flow within the shooting range area of the in-vehicle monocular camera to obtain a 3D scene after semantic segmentation.
[0043] S3. Perform integrated positioning and detour path generation using the 3D scene after semantic segmentation and the high-precision map. Specifically, match the 3D scene after semantic segmentation with the high-precision map to obtain the vehicle's real-time positioning information; then, obtain the drivable area corresponding to the vehicle according to the 3D scene after semantic segmentation and the positioning position of the current vehicle; finally, use the detour path algorithm combined with the drivable area corresponding to the vehicle to generate a detour path, where the detour path algorithm includes but is not limited to the A* algorithm, Dijkstra algorithm, and RRT (Rapidly-Exploring Random Tree) algorithm.
[0044] The working principle of the above technical solution is as follows: The on-vehicle monocular camera continuously captures the environment around the vehicle to obtain continuous video stream data. The monocular camera captures light through the lens and converts it into an electrical signal, which then undergoes a series of signal processing and conversions to form digital video stream data for subsequent analysis. This data contains various information around the vehicle, such as roads, vehicles, pedestrians, traffic signs, and traffic lights. Then, semantic segmentation is performed on the video stream data using the vehicle's driving speed characteristics, as well as the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the on-vehicle monocular camera to obtain the 3D scene after semantic segmentation. The 3D scene after semantic segmentation is fused with the high-precision map. The high-precision map contains detailed road information, such as the positions of lane lines, traffic signs, road gradients, and curvatures. By matching and comparing the objects and features in the 3D scene with the information on the high-precision map, the vehicle can determine its exact position on the map. For example, by identifying the lane lines in the video and matching them with the lane line information on the high-precision map, the vehicle can determine the lane it is currently in and its driving direction. After determining the vehicle's position, based on the surrounding traffic conditions (obtained from the 3D scene after semantic segmentation) and the destination information, a detour path is generated using a path planning algorithm. The path planning algorithm takes into account various factors, such as the traffic capacity of the road, traffic flow, speed limits, etc., to generate a safe and efficient driving path. If an obstacle or traffic congestion is detected in the 3D scene ahead, the algorithm will automatically find other feasible routes to guide the vehicle to detour and ensure that the vehicle can reach the destination smoothly.
[0045] The effects of the above technical solution are as follows: By combining the vehicle's driving speed characteristics, pedestrian flow change rate, and vehicle flow change rate to perform semantic segmentation on the video stream data, different objects and scene elements can be distinguished more accurately. For example, when the vehicle is driving at a high speed, distant traffic signs, pedestrians, and other targets can be identified more precisely; according to the pedestrian flow and vehicle flow change rates, dynamic objects, such as pedestrians crossing the road and vehicles changing lanes, can be distinguished more clearly, thereby improving the accuracy of semantic segmentation and providing a more reliable basis for subsequent positioning and path planning. Based on the accurate 3D scene after semantic segmentation and the high-precision map for comprehensive positioning and detour path generation, the vehicle's position on the map can be determined more precisely, reducing the positioning error. At the same time, when generating the detour path, the actual traffic conditions on the road, such as dynamic pedestrian and vehicle flows, can be fully considered to plan a safer and more efficient driving path, avoiding accidents or delays caused by inaccurate positioning or unreasonable path planning.
[0046] Meanwhile, the above technical solution processes the video stream data by using information such as the vehicle driving speed characteristics, the pedestrian flow change rate, and the vehicle flow change rate, and can quickly screen and analyze key information, improving the speed of semantic segmentation. This real-time processing ability enables the vehicle to respond promptly to changes in the road environment, generating accurate semantic segmentation results and 3D scenes in a short time, providing timely data support for subsequent positioning and path planning. Since the 3D scene after semantic segmentation can be obtained in real time, the vehicle can adjust the detour path in a timely manner according to the real-time changes in the road environment, such as suddenly appearing obstacles, traffic congestion, etc. This real-time adjustment ability improves the vehicle's ability to handle emergencies and ensures the safety and efficiency of autonomous driving.
[0047] In an embodiment of the present invention, semantic segmentation is performed on the video stream data by using the vehicle driving speed characteristics and the pedestrian flow change rate and the vehicle flow change rate within the shooting range area of the on-vehicle monocular camera to obtain a 3D scene after semantic segmentation, including:
[0048] S201. Process the video stream data by using the vehicle driving speed characteristics and the pedestrian flow change rate and the vehicle flow change rate within the shooting range area of the on-vehicle monocular camera to obtain a depth map, a material attribute map, and a 3D scene flow corresponding to the video stream data; wherein, the material attribute map maps the material feature vector m to 4 channels, and the index parameters corresponding to the 4 channels include diffuse reflectivity, specular reflectivity, transparency, and roughness;
[0049] S202. Perform semantic segmentation on the 3D scene flow to obtain an initial semantic segmentation result corresponding to the 3D scene flow;
[0050] S203. Extract and fuse features from the continuous RGB video frames, depth map, and material attribute map corresponding to the video stream data to obtain a multi-modal feature pyramid; wherein, the multi-modal feature pyramid contains color, geometric, and material information;
[0051] S204. Repair the initial semantic segmentation result by using the multi-modal feature pyramid and the 3D scene flow to obtain a repaired semantic segmentation result; wherein, the semantic segmentation result includes but is not limited to pedestrians, objects, roads, obstacles, puddles, traffic element lights, etc.;
[0052] S205. Annotate the 3D scene flow according to the semantic segmentation result to obtain a 3D scene after semantic segmentation.
[0053] The working principle of the above technical solution is as follows: By utilizing the vehicle driving speed characteristics and the change rates of the pedestrian flow and vehicle flow within the shooting range area of the on-vehicle monocular camera, the acquired video stream data is processed to generate a depth map, a material property map, and a 3D scene flow. The depth map can reflect the distance information between the objects in the scene and the camera. The material property map maps the material feature vectors to 4 channels (diffuse reflectance, specular reflectance, transparency, and roughness) to describe the material characteristics of the object surface, while the 3D scene flow reflects the motion information of the objects in the scene. The generation of these data is based on the comprehensive analysis and calculation of the pixel information in the video stream and information such as speed, pedestrian flow, and vehicle flow change rates. Semantic segmentation is performed on the generated 3D scene flow to obtain the initial semantic segmentation result corresponding to the 3D scene flow. The above technical solution uses a semantic segmentation algorithm to preliminarily classify different parts of the scene into different semantic categories, such as pedestrians and vehicles, according to information such as the motion characteristics of the objects in the 3D scene flow. Feature extraction is performed on the continuous RGB video frames, depth map, and material property map corresponding to the video stream data, and then the features from different modalities (color, geometry, and material) are fused to construct a multi-modal feature pyramid. The feature extraction process may use methods such as convolutional neural networks (CNNs) to extract representative features from different images and data. The multi-modal feature pyramid integrates information from multiple aspects such as color, geometry, and material, providing richer information support for the subsequent repair of the semantic segmentation result. Using the constructed multi-modal feature pyramid and the 3D scene flow, the previously obtained initial semantic segmentation result is repaired. By comprehensively considering the multi-modal features and the motion information in the scene flow, the errors, omissions, or inaccuracies that may exist in the initial result are corrected and improved to obtain a more accurate semantic segmentation result, which covers multiple semantic categories such as pedestrians, objects, roads, obstacles, puddles, and traffic elements. According to the repaired semantic segmentation result, the 3D scene flow is labeled, that is, each part is marked as the corresponding semantic category, so as to obtain the 3D scene after semantic segmentation. This 3D scene provides detailed and accurate environmental information for subsequent autonomous driving decisions.
[0054] The effects of the above technical solution are as follows: By combining information from multiple aspects such as vehicle driving speed, the change rates of pedestrian flow and vehicle flow, and comprehensively utilizing multi-modal data (RGB video frames, depth maps, material property maps, and 3D scene flows) for semantic segmentation and result repair, various elements in the scene can be identified and classified more accurately. For example, the analysis of material properties can help distinguish objects of different materials, and the multi-modal feature fusion provides richer information to determine the category of objects, thereby improving the accuracy of semantic segmentation and further providing more reliable environmental perception for autonomous driving. The use of multi-modal data enables the system to have stronger resistance to different environmental conditions and data noise. For example, in the case of light changes, occlusion, etc., information such as depth maps and material property maps can assist RGB video frames in accurate object recognition, and 3D scene flows can provide clues about object motion, helping the system maintain stable semantic segmentation performance in complex environments and improving the robustness of the system. Although processing multi-modal data may increase a certain amount of computational effort, through reasonable algorithm design and optimization, using information such as speed, pedestrian flow, and vehicle flow change rates can quickly screen and process key data. At the same time, the construction of a multi-modal feature pyramid can also improve the efficiency of feature extraction and fusion to a certain extent, thus ensuring that in autonomous driving scenarios with high real-time requirements, a 3D scene after semantic segmentation can be obtained in a timely and accurate manner to support the real-time decision-making of the vehicle. The semantic segmentation results cover a rich variety of semantic categories, including pedestrians, objects, roads, obstacles, puddles, and traffic elements, etc. And through the analysis of material properties and the fusion of multi-modal information, the system's understanding of the scene is deeper and more comprehensive. This helps the vehicle better perceive the surrounding environment and make more reasonable decisions, such as avoiding obstacles and choosing appropriate driving paths, etc., improving the safety and intelligence of autonomous driving.
[0055] In an embodiment of the present invention, the video stream data is processed by using the vehicle driving speed characteristics and the change rates of pedestrian flow and vehicle flow within the shooting range area of the in-vehicle monocular camera to obtain a depth map, a material property map, and a 3D scene flow corresponding to the video stream data, including:
[0056] S2011. Perform frame processing on the video stream data to obtain continuous RGB video frames corresponding to the video stream data;
[0057] S2012. Set the number of consecutive frames input to the dynamic neural radiance field NeRF model according to the maximum and minimum driving speeds of the vehicle within a preset time range in combination with the change rates of pedestrian flow and vehicle flow within the current shooting range area of the in-vehicle monocular camera.
[0058] S2013. Divide the consecutive RGB video frames according to the consecutive frame numbers, and input the divided consecutive RGB video frames into the dynamic neural radiance field NeRF model to output the depth map, material property map, and 3D scene flow corresponding to the video stream data.
[0059] The working principle of the above technical solution is as follows: perform frame processing on the video stream data collected by the vehicle-mounted monocular camera, and split the consecutive video stream into one frame of image after another to obtain the consecutive RGB video frames corresponding to the video stream data. These RGB video frames contain the visual information of the vehicle's surrounding environment and are the basic data for subsequent processing. First, monitor and record the maximum and minimum driving speeds of the vehicle within a preset time range in real time. The preset time range can be set specifically through experiments according to the actual performance parameters of the vehicle (such as the 0-100 km / h acceleration time, etc.) to ensure that it can effectively reflect the vehicle's speed change situation. At the same time, analyze each consecutive RGB video frame, identify the number of pedestrians and vehicles appearing in the frame through algorithms such as object detection, and then calculate the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the current vehicle-mounted monocular camera.
[0060] Then, based on the maximum speed, minimum speed, pedestrian flow change rate, and vehicle flow change rate, combined with the preset basic consecutive frame number (not less than 5), calculate and set the consecutive frame number input to the dynamic neural radiance field NeRF model through a specific formula. This process takes into account the vehicle speed state and the dynamic changes in the surrounding environment, and dynamically determines the appropriate input frame number.
[0061] Data input and output: Divide the previously obtained consecutive RGB video frames according to the determined consecutive frame numbers to form a video sequence that meets the requirements. Input these divided consecutive RGB video frames into the dynamic neural radiance field NeRF model. The NeRF model is a deep learning-based technology that can process the input video sequence. By learning and analyzing factors such as the visual information, vehicle speed, and environmental changes in the video frames, it outputs the depth map, material property map, and 3D scene flow corresponding to the video stream data. The depth map reflects the distance information between the objects in the scene and the camera, the material property map describes the material characteristics of the object surface (such as diffuse reflectivity, specular reflectivity, transparency, and roughness, etc.), and the 3D scene flow reflects the motion information of the objects in the scene. Specifically, the video processing process of the dynamic neural radiance field NeRF model includes:
[0062] Spatio-temporal modeling: Input the video frame sequence (including timestamps) into the NeRF network, and implicitly express the geometry (volume density) and material properties (reflectivity, transparency) of the 3D scene through the MLP.
[0063] Dynamic separation: Introduce an implicit deformation field to distinguish the motion trajectories of static backgrounds and dynamic objects.
[0064] Volume rendering: Generate the depth map, material property map, and 3D scene flow (3D motion vectors between adjacent frames) of the current view through ray casting.
[0065] Finally, output the depth map (object distance information), material property map (physical properties), and 3D scene flow (dynamic object motion information) that are aligned with the input video frames. The above technical solution introduces a video frame sequence (including timestamps) to provide physical property priors for subsequent segmentation, solving the problem of misjudgment of materials caused by traditional methods relying on pure RGB data.
[0066] The technical effects of the above technical solution are as follows: Dynamically adjust the number of consecutive frames input to the NeRF model according to the vehicle speed and the change rates of the pedestrian flow and vehicle flow, so that the model can better adapt to different driving scenarios. In complex and changeable scenarios (such as areas with high pedestrian flow, high vehicle flow, and large vehicle speed changes), increasing the number of input frames can provide more abundant information, enabling the model to more accurately generate the depth map, material property map, and 3D scene flow, thereby more precisely perceiving the surrounding environment, including the position, material, and motion state of objects, etc., providing a more reliable basis for autonomous driving decisions. This technical solution can dynamically adjust the input data according to the actual driving situation, making the NeRF model more adaptable to different road conditions, traffic conditions, and vehicle driving states. Whether on urban congested roads, highways, or rural roads, the model can obtain appropriate inputs through reasonable settings of the number of consecutive frames, accurately process the video stream data, output high-quality environmental representation information, and improve the generalization ability of the model in various scenarios. It avoids the problem of waste or insufficiency of computing resources that may be caused by fixed-frame number input. When the scene is relatively stable, the vehicle speed changes little, and the pedestrian and vehicle flow is low, reducing the number of input frames can reduce the computational load of the model, improve the operation efficiency, and save computing resources; while when more information is needed to process complex scenarios, increase the number of frames to ensure the accuracy of the model. This dynamic adjustment mechanism realizes the optimal allocation of computing resources, improves the real-time performance and resource utilization efficiency of the system without affecting the model performance. By combining the vehicle driving speed characteristics and the change rates of the pedestrian flow and vehicle flow to process the video stream data and input it into the NeRF model, the model can better capture the dynamic changes and detailed information in the scene. This helps to generate more accurate and complete depth maps, material property maps, and 3D scene flows, improve the quality of 3D scene reconstruction, provide a more realistic and accurate virtual environment representation for the autonomous driving system, and support more advanced autonomous driving functions, such as path planning, obstacle detection, and avoidance.
[0067] An embodiment of the present invention sets the number of consecutive frames input to the dynamic neural radiance field NeRF model according to the maximum and minimum driving speeds of the vehicle within a preset time range in combination with the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the current on-vehicle monocular camera, including:
[0068] Step 1: Identify the number of pedestrians and vehicles appearing in each consecutive RGB video frame by recognizing the consecutive RGB video frames;
[0069] Step 2: Obtain the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the current on-vehicle monocular camera according to the number of pedestrians and vehicles appearing in each consecutive RGB video frame;
[0070] Step 3: Real-time monitor the maximum speed and minimum speed of the current vehicle within a preset time range;
[0071] Step 4: Set the number of consecutive frames (monocular video sequence) input to the dynamic neural radiance field NeRF model according to the maximum and minimum speeds in combination with the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the current on-vehicle monocular camera;
[0072] Among them, the number of consecutive frames (monocular video sequence) input to the dynamic neural radiance field NeRF model is obtained through the following formula:
[0073]
[0074] Among them, N represents the number of consecutive frames input to the dynamic neural radiance field NeRF model, and N is rounded up; n represents the preset basic number of consecutive frames, and the value of the number of consecutive frames shall not be less than 5; B r represents the change rate of the pedestrian flow around the vehicle; B c represents the change rate of the vehicle flow around the vehicle; r represents the speed state parameter; among them, the speed state parameter is obtained through the following formula:
[0075]
[0076] Among them, r represents the speed state parameter; v max and v min respectively represent the maximum speed and minimum speed of the current vehicle within a preset time range; t represents the time length corresponding to the preset time range, and the value of the preset time range is set experimentally according to the actual performance parameters of the vehicle (for example, the 0-100 km / h acceleration time, etc.).
[0077] The working principle of the above technical solution is as follows: Identify the continuous RGB video frames obtained by the in-vehicle monocular camera, and determine the number of pedestrians and vehicles appearing in each continuous RGB video frame through computer vision-related algorithms (such as object detection algorithms). The above technical solution is the basis for subsequent calculation of the pedestrian flow change rate and the vehicle flow change rate. By detecting and counting the objects in the video frames, the distribution information of people and vehicles in the scene is obtained. According to the number of pedestrians and vehicles appearing in each continuous RGB video frame, calculate the pedestrian flow change rate and the vehicle flow change rate within the shooting range area of the current in-vehicle monocular camera. Real-time monitor the maximum speed and minimum speed of the current vehicle within a preset time range. The speed information of the vehicle can be obtained through devices such as the vehicle's own speed sensor. The length of the preset time range is set experimentally according to the actual performance parameters of the vehicle (such as the zero-to-hundred-kilometer acceleration time, etc.) to ensure that the obtained speed data can reflect the speed change characteristics of the vehicle during this time period.
[0078] The effect of the above technical solution is as follows: Dynamically set the number of consecutive frames input to the NeRF model according to the vehicle speed and the change rates of the surrounding pedestrian flow and vehicle flow, enabling the model to better adapt to different driving scenarios. In the case of large vehicle speed changes or frequent changes in the surrounding pedestrian and vehicle traffic, increasing the number of input frames can provide more abundant information to help the model more accurately reconstruct the scene and understand the environment; while when the scene is relatively stable, reducing the number of input frames can reduce the computational amount and improve the running efficiency of the model. Reasonable setting of the number of consecutive frames helps the NeRF model to more comprehensively capture the dynamic information in the scene. For example, in areas with large pedestrian and vehicle traffic, a larger number of consecutive frames can enable the model to better track the movement trajectories of pedestrians and vehicles, thus more accurately modeling and analyzing the scene, enhancing the model's understanding and processing capabilities for complex scenes. It avoids the problem of waste or insufficiency of computational resources that may be caused by fixed-frame input. By dynamically adjusting the number of consecutive frames according to the actual situation, while ensuring the model performance, the computational resources are effectively utilized. When the vehicle speed is stable and the surrounding environment changes little, reducing the number of input frames can reduce the computational cost and improve the real-time performance of the system; while when more information is needed to process complex scenes, increasing the number of frames can ensure the accuracy of the model, realizing the optimal allocation of computational resources. Combining the vehicle speed and the environmental change rate to set the number of consecutive frames provides more appropriate input data for the NeRF model, which helps to improve the training and inference accuracy of the model. More accurate scene modeling and environmental understanding can provide more reliable environmental perception for the autonomous driving system, thus supporting safer and more intelligent driving decisions and enhancing the overall performance of autonomous driving.
[0079] In one embodiment of the present invention, feature extraction and fusion are performed on the continuous RGB video frames, depth maps, and material property maps corresponding to the video stream data to obtain a multi-modal feature pyramid, including:
[0080] S2031. Respectively use a convolutional neural network to generate semantic feature maps corresponding to consecutive RGB video frames, material property maps, and depth maps;
[0081] S2032. Using the scale of the semantic feature map of the consecutive RGB video as a standard, adjust the semantic feature maps corresponding to the material property map and the depth map to be consistent with the scale of the RGB video semantic feature map, and project the adjusted semantic feature maps corresponding to the material property map and the depth map into the semantic space matching the RGB features for weighted fusion to obtain a multi-modal feature pyramid.
[0082] Specifically: Use a convolutional neural network to extract features from consecutive RGB video frames to obtain multi-scale RGB semantic feature maps; use a convolutional neural network to extract features from the material property map to obtain the semantic feature map corresponding to the material property map; use a convolutional neural network to extract features from the depth map to obtain the depth feature map corresponding to the depth map; use bilinear interpolation to adjust the resolutions of the semantic feature map corresponding to the material property map and the depth feature map to be consistent with the RGB semantic feature map; project the semantic feature map corresponding to the material property map and the depth feature map into the semantic space matching the RGB features for weighted fusion to obtain a multi-modal feature pyramid.
[0083] The working principle of the above technical solution is to use a convolutional neural network (CNN) to process consecutive RGB video frames. The CNN automatically extracts various semantic features in the video frames, such as edges, textures, shapes, etc., through structures such as convolutional layers and pooling layers, to generate multi-scale RGB semantic feature maps. Feature maps of different scales contain different levels of semantic information. Small-scale feature maps may capture more detailed information, while large-scale feature maps reflect more macroscopic scene structures. Similarly, use the CNN to extract features from the material property map. The material property map contains information where the material feature vector is mapped to 4 channels (diffuse reflectance, specular reflectance, transparency, and roughness). The CNN extracts semantic features related to the material properties from these channels to obtain the semantic feature map corresponding to the material property map. Use the CNN to extract features from the depth map. The depth map represents the distance information between the objects in the scene and the camera. The CNN can extract semantic features related to the spatial structure, object positions, etc. from the depth data to generate the depth feature map corresponding to the depth map. Using the scale of the semantic feature map of the consecutive RGB video as a standard, use the bilinear interpolation method to process the semantic feature map corresponding to the material property map and the depth feature map. Bilinear interpolation is a commonly used image scaling technique. By calculating the weighted average between adjacent pixels, the resolutions of the semantic feature maps of the material property map and the depth map are adjusted to be the same as the RGB semantic feature Figure 1To make the feature maps of different modalities comparable in scale. Map the semantic feature map and depth feature map corresponding to the material property map with adjusted scale to the semantic space matching the RGB features. This step is to fuse the features of different modalities at the same semantic level. Then, perform weighted fusion on these mapped feature maps, assign corresponding weights according to the importance of different modality features, and combine them together to form a multi-modal feature pyramid. The multi-modal feature pyramid integrates information from RGB video frames, material property maps, and depth maps, contains multi-faceted semantic features such as color, geometry, and materials, and provides richer information for subsequent analysis and processing.
[0084] The effects of the above technical solution are as follows: By fusing multi-modal features from consecutive RGB video frames, material property maps, and depth maps, the multi-modal feature pyramid can represent scene information more comprehensively. It not only contains visual color and texture information (from RGB video frames), but also covers the spatial position information of objects (from depth maps) and material property information (from material property maps). This rich feature representation helps to more accurately identify and understand objects and elements in the scene, improving the system's perception ability of complex scenes. Different modality data have different stabilities in the face of various environmental factors (such as lighting changes, occlusions, etc.). For example, depth maps can still provide good object position information in low-light conditions, while RGB video frames have advantages in color and texture recognition. Through multi-modal feature fusion, the system can comprehensively utilize the advantages of different modalities, reduce the limitations of single-modal data, enhance the resistance to environmental changes, and improve the robustness and reliability of the system. The multi-modal feature pyramid provides richer and more accurate feature inputs for subsequent tasks (such as semantic segmentation, object detection, etc.). In semantic segmentation, combining multi-modal features can more precisely classify each pixel and distinguish different object and scene regions; in object detection, it helps to more accurately locate and identify target objects, reducing the false detection and missed detection rates. Therefore, this technical solution can significantly improve the accuracy and performance of related tasks. Since multi-modal feature fusion can capture more comprehensive and diverse information in the scene, the model trained based on the multi-modal feature pyramid has stronger generalization ability when facing different scenes and datasets. The model can better adapt to different environmental conditions, object types, and scene layouts, and can also perform well in new and unseen scenes, improving the application scope and practicality of the model.
[0085] An embodiment of the present invention uses a multi-modal feature pyramid and 3D scene flow to repair the initial result of semantic segmentation and obtain the repaired semantic segmentation result, including:
[0086] S2041. Extract features from the 3D voxel grid generated from consecutive RGB video frames using a 3D sparse convolutional kernel to obtain spatio-temporal features. Specifically: Stack the 2D feature maps in the consecutive RGB video frames corresponding to the consecutive frame numbers into a 3D voxel grid; Slide the 3D sparse convolutional kernel over the 3D voxel grid to extract spatio-temporal features.
[0087] S2042. Calculate the occlusion region mask of the current frame in the consecutive RGB video frames by combining the current driving speed of the vehicle, the change rate of the pedestrian flow around the vehicle, and the change rate of the traffic flow with the 3D scene flow and its corresponding density map, and obtain the occlusion regions that appear in the current frame.
[0088] S2043. Repair the occlusion regions to generate a repaired semantic segmentation result. Specifically: Use the historical frame features in the consecutive RGB video frames through the LSTM memory unit to repair the occlusion regions of the current frame to obtain the repaired spatio-temporal features; Then, upsample the repaired spatio-temporal features to generate a repaired semantic segmentation result.
[0089] The working principle of the above technical solution is as follows: First, stack the 2D feature maps in consecutive RGB video frames corresponding to consecutive frame numbers into a 3D voxel grid. This step combines the 2D information of multiple consecutive frames into a 3D structure with spatial and temporal dimensions for subsequent extraction of spatio-temporal features. Then, use a 3D sparse convolutional kernel to perform a sliding operation on the 3D voxel grid. The 3D sparse convolutional kernel moves on the 3D voxel grid and extracts the spatio-temporal features therein through convolutional operations. This operation can capture the motion information of objects in the video frame sequence and the spatial relationships between different frames, thereby obtaining a feature representation containing information in both time and space dimensions. Combine the current driving speed of the vehicle, the change rate of the pedestrian flow around the vehicle, and the change rate of the traffic flow, and at the same time refer to the 3D scene flow and its corresponding density map. These information comprehensively reflect the motion state of the vehicle and the dynamic changes in the surrounding environment. Through the analysis and calculation of this information, determine the occlusion area mask of the current frame in the consecutive RGB video frames. The occlusion area mask is an identifier indicating which areas in the current frame are occluded. In this way, the occluded areas that appear in the current frame can be accurately located. Use an LSTM (Long Short-Term Memory Network) memory unit to repair the occluded areas of the current frame using the features of historical frames in the consecutive RGB video frames. The LSTM has a memory function and can remember the relevant information in the historical frames and apply it to the repair of the occluded areas of the current frame. In this way, the features of the unoccluded parts in the historical frames can be used to infer and fill the occluded areas of the current frame, thereby obtaining the repaired spatio-temporal features. Perform an upsampling operation on the repaired spatio-temporal features. Upsampling is the process of converting a low-resolution feature map into a high-resolution one. Through upsampling, the resolution of the repaired feature map can be made to match that of the original semantic segmentation result, and finally generate the repaired semantic segmentation result.
[0090] The effects of the above technical solution are as follows: By extracting spatio-temporal features and repairing occluded regions, objects and regions in the scene can be identified and segmented more accurately. The extraction of spatio-temporal features can capture the motion information of objects, which helps to track and segment objects more accurately in dynamic scenes; the repair of occluded regions solves the problem of inaccurate segmentation caused by occlusion, making the semantic segmentation results more complete and accurate, thus improving the overall semantic segmentation accuracy. This technical solution comprehensively considers various factors such as vehicle driving speed, pedestrian flow, and traffic flow change rate, as well as information such as 3D scene flow and density maps, enabling the system to have stronger adaptability to complex and changing environments. In different traffic scenes and environmental conditions, it can handle occlusion situations and object motion changes more accurately, reduce segmentation errors caused by environmental factors, and improve the robustness and reliability of the system. The extraction of spatio-temporal features and the use of LSTM to remember and apply historical frame features enable the system to better handle dynamic scenes. It can accurately track the motion trajectory of objects in consecutive frames. Even when an object is occluded or partially visible, reasonable inferences and repairs can be made through historical information and spatio-temporal features, thereby improving the understanding and processing ability of dynamic scenes. The repair of occluded regions ensures the integrity of information in the semantic segmentation results. In actual scenes, occlusion is a common phenomenon. By repairing occluded regions, the information lost due to occlusion is avoided, enabling the semantic segmentation results to more comprehensively reflect the true situation of the scene and providing richer and more accurate information support for subsequent tasks such as autonomous driving decision-making.
[0091] In an embodiment of the present invention, the occluded region mask is 1 or 0. Wherein, when the occluded region mask is 1, it indicates that the pixel corresponding to the occluded region mask is marked as an occluded region; when the occluded region mask is 0, it indicates that the pixel corresponding to the occluded region mask is marked as a non-occluded region. And, the occluded region mask is obtained by combining the difference between the density value of pixel x in the previous consecutive RGB video frame and the density value of pixel x in the current consecutive RGB video frame with a dynamically adjusted occlusion threshold and a dynamically adjusted motion residual threshold.
[0092] Among them, the model of the occluded region mask is as follows:
[0093]
[0094] Among them, R t (p) represents the occluded region mask; Δδ(p) represents the difference between the density value of pixel x in the previous consecutive RGB video frame and the density value of pixel x in the current consecutive RGB video frame; wherein, the density value is the object occupancy probability, and Δδ(p) being greater than zero indicates that the object may disappear (i.e., be transparent) or be occluded and reflect light. Here, the expression "occupancy" represents the probability of being occluded, reflected, or transparent at this pixel point in specific practices; μ01 Represents the dynamically adjusted occlusion threshold; W r (x) represents the motion residual during vehicle driving, used to evaluate the difference between the actual observed density and the motion predicted density, and exclude misjudgments caused by rapid object movement; μ 02 Represents the dynamically adjusted motion residual threshold;
[0095] Among them, the expression of the motion residual during vehicle driving is as follows:
[0096] W r (x) = w t (x) - w s (x)
[0097] Among them, w s (x) represents the density value of pixel x in the current consecutive RGB video frame; w t (x) represents the density value after warping the previous frame density map to the current frame based on the scene flow.
[0098] Among them, the dynamically adjusted occlusion threshold is obtained through the following formula:
[0099]
[0100] Among them, μ 01 Represents the dynamically adjusted occlusion threshold; μ z0 Represents the preset reference occlusion threshold, preferably 0.5; v represents the current driving speed of the vehicle; v cy Represents the preset speed reference value; B r Represents the change rate of the number of people around the vehicle; B c Represents the change rate of the number of vehicles around the vehicle;
[0101] And, the dynamically adjusted motion residual threshold is obtained through the following formula:
[0102]
[0103] Among them, μ 02 Represents the dynamically adjusted motion residual threshold; μ c0 Represents the preset reference motion residual threshold, preferably 0.3; v represents the current driving speed of the vehicle; v cy Represents the preset speed reference value.
[0104] The working principle and effects of the above technical solution are as follows: By comprehensively considering the density difference, motion residual, and thresholds dynamically adjusted according to vehicle speed, pedestrian flow, and traffic flow, the above technical solution can more accurately detect occluded areas. Dynamically adjusting the thresholds enables the system to adaptively judge occlusion situations according to different driving scenarios (such as vehicle speed, pedestrian flow, and traffic flow), avoiding misjudgments or missed detections that may be caused by fixed thresholds, thereby improving the accuracy of occlusion detection. The introduction of motion residual effectively eliminates misjudgments caused by the rapid movement of objects. During vehicle driving, the rapid movement of objects may cause changes in density values, which are easily misjudged as occlusion situations. By calculating the motion residual and comparing it with the dynamically adjusted motion residual threshold, it is possible to accurately distinguish between true occlusions and normal object movements, enhancing the system's resistance to interference factors and improving the system's stability. Dynamically adjusting the occlusion threshold and motion residual threshold according to the change rates of vehicle speed, pedestrian flow, and traffic flow enables the system to better adapt to complex and changing traffic scenarios. In different scenarios, such as urban congested sections and highways, the system can adjust the thresholds according to the actual situation to accurately detect occluded areas, improving the system's adaptability and reliability in various complex scenarios. Accurate detection of occluded areas provides a more reliable basis for subsequent semantic segmentation. During semantic segmentation, occluded areas can be processed more accurately, avoiding segmentation errors caused by occlusion, thereby improving the accuracy of semantic segmentation and making the segmentation results better reflect real scene information, which helps the autonomous driving system make more reasonable decisions.
[0105] At the same time, the above technical solution, during semantic segmentation, not only relies on the video stream of the vehicle environment but also considers the operating state during vehicle driving and the actual states of road pedestrians and vehicles for semantic segmentation. It can effectively improve the accuracy of semantic segmentation while greatly enhancing the followability of semantic segmentation with vehicle operation and the automatic adjustability of semantic segmentation, thereby maximizing the sensitivity and flexibility of semantic segmentation. At the same time, the above technical solution can effectively improve the matching between semantic segmentation and vehicle driving state and road state. Furthermore, when the body is temporarily occluded (such as a pedestrian reappearing after being briefly occluded by a vehicle), it can improve the accuracy and processing efficiency of semantic segmentation. Also, in the case of rare object materials (such as transparent glass, reflective road surfaces) or special lighting (such as strong specular reflection) on the road, it can improve the accuracy and processing efficiency of semantic segmentation. At the same time, because the operating state during vehicle driving and the actual states of road pedestrians and vehicles are utilized during semantic segmentation, it can adjust the mobility and sensitivity of semantic segmentation to handle the above special situations according to the operating state during vehicle driving and the actual states of road pedestrians and vehicles. Furthermore, on the premise of ensuring the adaptation of vehicle operation state and semantic segmentation operation rate, while improving the accuracy of intelligent recognition, it can maximize the release of in-vehicle system computing power and reduce the computing power load of the in-vehicle system.
[0106] An embodiment of the present invention provides an autonomous driving system based on machine vision, such as Figure 2 shown, the autonomous driving system based on machine vision includes:
[0107] A video stream data acquisition module, configured to acquire video stream data around the vehicle by using an in-vehicle monocular camera; specifically, adjust the shooting parameters of the in-vehicle monocular camera in real time according to the environmental conditions where the vehicle is located, where the environmental conditions include but are not limited to light intensity and weather conditions; the shooting parameters include but are not limited to resolution, field of view angle, and dynamic range; then, after the shooting parameters of the in-vehicle monocular camera are adjusted, control the in-vehicle monocular camera to perform real-time shooting in real time to acquire video stream data around the vehicle.
[0108] A semantic segmentation module, configured to perform semantic segmentation on the video stream data by using the vehicle driving speed characteristics and the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the in-vehicle monocular camera to obtain a 3D scene after semantic segmentation;
[0109] A path generation module, configured to perform integrated positioning and detour path generation by using the 3D scene after semantic segmentation and a high-precision map. Specifically, match the 3D scene after semantic segmentation with the high-precision map to obtain the vehicle's real-time positioning information; then, obtain the drivable area corresponding to the vehicle according to the 3D scene after semantic segmentation and the current vehicle's positioning position; finally, generate a detour path by using a detour path algorithm in combination with the drivable area corresponding to the vehicle, where the detour path algorithm includes but is not limited to the A* algorithm, Dijkstra algorithm, and RRT (Rapidly-Exploring Random Tree) algorithm.
[0110] The working principle of the above technical solution is as follows: The on-vehicle monocular camera continuously captures the environment around the vehicle to obtain continuous video stream data. The monocular camera captures light through the lens and converts it into an electrical signal, which then undergoes a series of signal processing and conversions to form digital video stream data for subsequent analysis. This data contains various information around the vehicle, such as roads, vehicles, pedestrians, traffic signs, and traffic lights. Then, semantic segmentation is performed on the video stream data using the vehicle's driving speed characteristics and the pedestrian flow change rate and vehicle flow change rate within the shooting range area of the on-vehicle monocular camera to obtain a 3D scene after semantic segmentation. The 3D scene after semantic segmentation is fused with the high-precision map. The high-precision map contains detailed road information, such as the positions of lane lines, traffic sign positions, road gradients, and curvatures. By matching and comparing the objects and features in the 3D scene with the information on the high-precision map, the vehicle can determine its exact position on the map. For example, by identifying the lane lines in the video and matching them with the lane line information on the high-precision map, the vehicle can determine the lane it is currently in and its driving direction. After determining the vehicle's position, based on the surrounding traffic conditions (obtained from the 3D scene after semantic segmentation) and the destination information, a detour path is generated using a path planning algorithm. The path planning algorithm takes into account various factors, such as the traffic capacity of the road, traffic flow, speed limits, etc., to generate a safe and efficient driving path. If an obstacle or traffic congestion is detected in the 3D scene ahead, the algorithm will automatically find other feasible routes to guide the vehicle to detour and ensure that the vehicle can reach the destination smoothly.
[0111] The effects of the above technical solution are as follows: By combining the vehicle's driving speed characteristics, pedestrian flow change rate, and vehicle flow change rate to perform semantic segmentation on the video stream data, different objects and scene elements can be distinguished more accurately. For example, when the vehicle is driving at a high speed, distant traffic signs, pedestrians, and other targets can be identified more precisely; according to the pedestrian flow and vehicle flow change rates, dynamic objects, such as pedestrians crossing the road and vehicles changing lanes, can be distinguished more clearly, thereby improving the accuracy of semantic segmentation and providing a more reliable basis for subsequent positioning and path planning. Based on the accurate 3D scene after semantic segmentation and the high-precision map for integrated positioning and detour path generation, the vehicle's position on the map can be determined more precisely, reducing positioning errors. At the same time, when generating a detour path, the actual traffic conditions on the road, such as dynamic pedestrian and vehicle flows, can be fully considered to plan a safer and more efficient driving path, avoiding accidents or delays caused by inaccurate positioning or unreasonable path planning.
[0112] Meanwhile, the above technical solution processes video stream data by using information such as vehicle driving speed characteristics, pedestrian flow change rate, and vehicle flow change rate, and can quickly screen and analyze key information, improving the speed of semantic segmentation. This real-time processing ability enables the vehicle to respond promptly to changes in the road environment, generate accurate semantic segmentation results and 3D scenes in a short time, and provide timely data support for subsequent positioning and path planning. Since the 3D scene after semantic segmentation can be obtained in real time, the vehicle can adjust the detour path in a timely manner according to the real-time changes in the road environment, such as suddenly appearing obstacles, traffic congestion, etc. This real-time adjustment ability improves the vehicle's ability to cope with emergencies and ensures the safety and efficiency of autonomous driving.
[0113] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and its equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. An automatic driving method based on machine vision, characterized in that: The machine vision-based autonomous driving method includes: Use the on-board monocular camera to obtain video stream data around the vehicle; The video stream data is semantically segmented using the vehicle's speed characteristics and the change rates of pedestrian and vehicle flows within the shooting range of the on-board monocular camera to obtain a 3D scene after semantic segmentation. The semantically segmented 3D scene and high-precision map are used for comprehensive positioning and detour path generation.
2. The automatic driving method based on machine vision according to claim 1, characterized in that: The method of obtaining video stream data around the vehicle by using a vehicle-mounted monocular camera includes: Adjust the shooting parameters of the on-board monocular camera in real time according to the vehicle's environmental conditions; After the shooting parameters of the on-board monocular camera are adjusted, the on-board monocular camera is controlled in real time to perform real-time shooting and obtain video stream data around the vehicle.
3. The automatic driving method based on machine vision according to claim 1, characterized in that: The video stream data is semantically segmented using the vehicle speed characteristics and the change rates of pedestrian and vehicle flows within the shooting range of the on-board monocular camera to obtain the 3D scene after semantic segmentation, including: The video stream data is processed using the vehicle's driving speed characteristics and the change rates of pedestrian flow and vehicle flow in the shooting range of the on-board monocular camera to obtain a depth map, a material property map, and a 3D scene flow corresponding to the video stream data; Performing semantic segmentation on the 3D scene stream to obtain an initial semantic segmentation result corresponding to the 3D scene stream; Perform feature extraction and fusion on the continuous RGB video frames, depth maps and material property maps corresponding to the video stream data to obtain a multimodal feature pyramid; The initial semantic segmentation result is repaired using the multimodal feature pyramid and 3D scene flow to obtain the repaired semantic segmentation result; The 3D scene stream is labeled according to the semantic segmentation result to obtain the 3D scene after semantic segmentation.
4. The automatic driving method based on machine vision according to claim 3, characterized in that: The video stream data is processed using the vehicle speed characteristics and the change rates of pedestrian flow and vehicle flow in the shooting range of the on-board monocular camera to obtain a depth map, a material property map and a 3D scene flow corresponding to the video stream data, including: Performing frame processing on the video stream data to obtain continuous RGB video frames corresponding to the video stream data; The number of continuous frames input to the dynamic neural radiation field NeRF model is set according to the maximum and minimum values of the vehicle's driving speed within a preset time range combined with the change rate of the pedestrian flow and the change rate of the vehicle flow within the shooting range of the current on-board monocular camera; The continuous RGB video frames are divided according to the number of consecutive frames, and the divided continuous RGB video frames are input into the dynamic neural radiation field NeRF model, and the depth map, material property map and 3D scene stream corresponding to the video stream data are output.
5. The automatic driving method based on machine vision according to claim 3, characterized in that: The continuous RGB video frames, depth maps and material property maps corresponding to the video stream data are extracted and fused to obtain a multimodal feature pyramid, including: Convolutional neural networks are used to generate semantic feature maps corresponding to continuous RGB video frames, material attribute maps, and depth maps; Taking the scale of the continuous RGB video semantic feature map as the standard, the semantic feature maps corresponding to the material attribute map and the depth map are adjusted to be consistent with the scale of the RGB video semantic feature map, and the adjusted semantic feature maps corresponding to the material attribute map and the depth map are projected into the semantic space matching the RGB feature for weighted fusion to obtain a multimodal feature pyramid.
6. The automatic driving method based on machine vision according to claim 3, characterized in that: The initial semantic segmentation result is repaired using the multimodal feature pyramid and 3D scene flow to obtain the repaired semantic segmentation result, including: The 3D voxel grid generated by continuous RGB video frames is extracted using 3D sparse convolution kernels to obtain spatiotemporal features; The occlusion area mask of the current frame in the continuous RGB video frames is calculated by using the current speed of the vehicle, the change rate of the flow of people and vehicles around the vehicle, and the 3D scene flow and its corresponding density map to obtain the occlusion area appearing in the current frame; The occluded area is repaired to generate the repaired semantic segmentation result.
7. The automatic driving method based on machine vision according to claim 6, characterized in that: The occlusion area mask is 1 or 0, wherein when the occlusion area mask is 1, it indicates that the pixel point corresponding to the occlusion area mask is marked as an occlusion area; when the occlusion area mask is 0, it indicates that the pixel point corresponding to the occlusion area mask is marked as a non-occlusion area.
8. The machine vision-based automatic driving method according to claim 6 or 7, characterized in that: The occlusion area mask is obtained by using the difference between the density value of pixel x in the previous continuous RGB video frame and the density value of pixel x in the current continuous RGB video frame combined with the dynamically adjusted occlusion threshold and the dynamically adjusted motion residual threshold.
9. An automatic driving system based on machine vision, characterized in that: The machine vision-based autonomous driving system includes: A video stream data acquisition module is used to acquire video stream data around the vehicle using a vehicle-mounted monocular camera; The semantic segmentation module is used to perform semantic segmentation on the video stream data by using the vehicle speed characteristics and the change rate of the pedestrian flow and the change rate of the vehicle flow within the shooting range of the on-board monocular camera, and obtain the 3D scene after semantic segmentation; The path generation module is used to perform comprehensive positioning and detour path generation using the semantically segmented 3D scene and high-precision map.
Citation Information
Patent Citations
Automatic driving video semantic segmentation system and method
CN112364822A
Driving path planning method
CN114119896A
Control strategy determination method and device, computer equipment, storage medium and product
CN118015586A
Lane line-based intelligent driving control method and apparatus, and electronic device
US20200293797A1