Multi-Sensor Fusion Collaborative Autonomous Driving Decision-Making Method for Complex Road Conditions

Through multi-sensor fusion collaborative autonomous driving decision-making method, identifying key information of unstructured roads and building virtual lane lines, solving the problems of perception and decision-making in complex road conditions and unstructured road environments in existing technologies, achieving higher environmental perception accuracy and safe driving.

CN119810784BActive Publication Date: 2025-05-30ZHEJIANG ELECTROMECHANICAL VOCATIONAL & TECH COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510246592.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-30
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

Existing autonomous driving technology is difficult to achieve effective environmental perception and safe driving decisions in complex road conditions and unstructured road environments, mainly due to its strong dependence on high-precision maps and structured roads.

Method used

Multi-sensor fusion collaborative autonomous driving decision-making method is adopted to collect road scene images through cameras, identify construction markers and passable areas, build virtual lane lines, and combine infrared light images to construct autonomous driving decisions.

Benefits of technology

Improves the accuracy and robustness of environmental perception, and can generate driving strategies suitable for complex road conditions in unstructured road environments to ensure the safe driving of the vehicle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810784B_ABST
    Figure CN119810784B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions, which relates to the technical field of image processing. The multi-sensor fusion collaborative autonomous driving decision-making method of the present invention specifically includes: collecting road scene images through a camera and identifying target objects in the front area to determine entry into a temporary construction area; dividing the passable area based on the road scene images; constructing virtual lane lines based on the passable area and the detected vehicle position information; constructing an autonomous driving decision according to the virtual lane lines in combination with the passable area and area targets. Through the above steps, the present invention improves the accuracy and robustness of environmental perception by collecting the environmental information around the vehicle in real time and using image vision processing algorithms to identify the key information of unstructured roads.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and particularly to a multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions, which is especially applicable to autonomous driving systems in unstructured road environments, aiming to solve the perception and decision-making problems of existing autonomous driving technologies in complex road conditions. Background Art

[0002] With the support of high-precision maps, current autonomous driving systems have been able to achieve many complex driving tasks, such as automatic lane change, path navigation, traffic signal recognition and other functions. These technologies are particularly mature for structured roads, such as urban arterial roads and highways. In these scenarios, high-precision maps can provide detailed environmental information, including the position of lane lines, road signs, traffic rules and the precise position of surrounding buildings, etc., thus providing reliable prior information for the environmental perception and path planning of autonomous driving vehicles. With the help of high-precision maps, autonomous driving systems can accurately determine the position of the vehicle and the obstacle information around it, and even be able to perceive future road condition changes in advance, so as to make efficient and safe driving decisions. Therefore, with the dual support of high-precision maps and structured roads, the application of autonomous driving technology has been promoted globally, especially the commercial implementation in some limited areas has achieved phased results, such as driverless freight vehicles on highways and driverless shuttle buses in closed parks.

[0003] However, although the above technologies perform well under specific conditions, the application scope of current autonomous driving technology is still severely limited. The core problem is that existing technologies rely strongly on high-precision maps and structured roads, and the construction and maintenance costs of these conditions are extremely high. First of all, the generation process of high-precision maps is very complex, which requires collecting a large amount of environmental data through professional equipment and accurately modeling with the help of efficient algorithms. This process not only requires a large amount of human and material resources, but also needs to be continuously updated to ensure the timeliness of the data. For example, the construction of urban roads, the opening of new roads, and the adjustment of traffic signs will all affect the accuracy of map data, and these dynamic changes are often difficult to be reflected in the map in a timely manner. In addition, the construction and maintenance of structured roads also require long-term investment in traffic infrastructure, such as clear road markings, a complete traffic signal system and standardized road design, etc. However, the real road environment is complex and diverse, especially in some rural, suburban or underdeveloped areas, the roads are not fully structured, and there may even be roads without markings, narrow or with complex obstacles. The lack of high-precision map support in these unstructured road environments makes it difficult for existing autonomous driving systems to adapt.

[0004] Meanwhile, the complexity of unstructured roads is also reflected in the dynamics of traffic changes. On the one hand, traffic construction in cities and rural areas changes every day. Factors such as temporary road closures, adjustments to construction areas, and newly added fork roads pose huge challenges to the update and maintenance of high-precision maps. On the other hand, unstructured roads such as rural paths, dirt roads, and gravel roads often do not have clear lane markings and standardized traffic signs, making it difficult for autonomous driving systems to obtain sufficient environmental perception information to complete safe driving. In addition, temporary obstacles pose higher requirements for the real-time perception and decision-making of autonomous driving systems. In this case, the defects of the existing technology are summarized as follows: on the one hand, autonomous driving systems lacking high-precision map support are prone to perception blind spots; on the other hand, in unstructured roads and complex dynamic road conditions, it is difficult to guarantee the decision-making reliability of autonomous driving systems. These problems severely limit the implementation and application of autonomous driving technology in a wider range of scenarios.

[0005] In view of the above problems, the present invention proposes a multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions. The present invention improves the accuracy and robustness of environmental perception by collecting environmental information around the vehicle in real time and using image vision processing algorithms to identify key information of unstructured roads. Based on this environmental model, a decision-making algorithm is used to generate driving strategies suitable for complex road conditions to ensure the safe driving of vehicles in unstructured road environments. Summary of the Invention

[0006] The present invention provides a multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions, which specifically includes the following steps:

[0007] S1: Collect road scene images through a camera and identify construction markers to determine entry into a temporary construction area;

[0008] S2: Divide the passable area based on the road scene image;

[0009] S3: Construct a virtual lane line based on the passable area and extract vehicle driving trajectory feature points;

[0010] S4: Construct an autonomous driving decision based on the virtual lane line and infrared light image.

[0011] The present invention provides a multi-sensor fusion collaborative autonomous driving decision-making system for new energy vehicles for complex road conditions, and the system includes:

[0012] Image acquisition device: Collect road scene images through a camera and identify target objects in the front area to determine entry into a temporary construction area;

[0013] Semantic segmentation module: The semantic segmentation module divides the passable area based on the road scene image; Virtual lane line generation module: The virtual lane line generation module constructs virtual lane lines based on the passable area and the detected vehicle position information;

[0014] Autopilot decision-making module: The autopilot decision-making module constructs an autopilot decision based on the virtual lane lines in combination with the passable area and regional targets.

[0015] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned multi-sensor fusion collaborative autopilot decision-making method for complex road conditions.

[0016] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned multi-sensor fusion collaborative autopilot decision-making method for complex road conditions.

[0017] Compared with the prior art, the present invention aims to solve the problem of driving decision-making for unstructured roads in the prior art. The present invention designs a multi-branch feature extraction network for the temporary construction scenario of unstructured roads, combines the characteristics of target features existing in the scenario, extracts features for multiple targets, and at the same time, in the feature enhancement part, introduces enhancement and suppression processing of foreground and background features to improve the detection accuracy of targets. In addition, the present invention obtains the passable area based on the detected targets to guide the semantic segmentation model, and constructs virtual lanes and virtual lane lines based on the vehicle occlusion situation. The detection of virtual lane lines makes full use of the position relationship in the current road environment, and solves the problem of autopilot decision-making in the temporary traffic environment on unstructured roads. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is the multi-sensor fusion collaborative autopilot decision-making for new energy vehicles facing complex road conditions in the present application. Detailed Embodiments

[0020] The embodiments of the present application will be described in detail below in conjunction with the drawings.

[0021] The following describes the implementation manners of the present application through specific examples. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope protected by the present application.

[0022] The embodiment of this specification proposes a multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions, as Figure 1 shown. The method specifically includes the following steps:

[0023] S1: Collect road scene images through a camera, identify the targets in the front area, and determine entering a temporary construction area; S2: Divide the passable area based on the road scene images;

[0024] S3: Construct a virtual lane line based on the passable area and the detected vehicle position information

[0025] S4: Construct an autonomous driving decision according to the virtual lane line in combination with the passable area and the area targets.

[0026] During the implementation process, the system adopts a collaborative design of the front-mounted camera above the vehicle head and the side cameras, ensuring comprehensive coverage of complex road conditions. Especially in an unstructured road environment, it can efficiently capture the key features of the road. To ensure complete road scene coverage, the front-mounted camera above the vehicle head is installed at the central position of the upper edge of the front windshield of the new energy vehicle, and the side cameras are respectively installed near the rearview mirrors on both sides of the vehicle. The combined cameras ensure effective coverage of the road directly in front of the vehicle by the front camera, and at the same time expand the vehicle's lateral perception ability through the side cameras. Especially in unstructured scenarios such as narrow roads, potential obstacles and special road condition information on both sides of the vehicle can be captured. The front camera has a horizontal viewing angle of 120° and a vertical viewing angle of 70°, while the horizontal viewing angle of the side camera is designed to be 90° and the vertical viewing angle is 60°, so as to achieve seamless connection of the perception ranges of multiple cameras and avoid the appearance of perception blind spots. To ensure the quality of the collected images, the resolution of the cameras is set to 1920×1080 pixels, and the frame rate is 30 frames per second, which can continuously capture clear road scene images when the vehicle is driving at high speed.

[0027] During the data acquisition process, the image data from the front camera and the side camera are transmitted to the in-vehicle computing platform through a high-speed bus. To ensure the synchronization of multi-camera data, the system adopts timestamp synchronization technology to correspond each frame of image with the vehicle's position information and motion state, ensuring accurate alignment during multi-sensor data fusion. In addition, to reduce transmission latency, the system introduces the H.265 video compression algorithm to perform lossless compression on the image data, saving bandwidth while maintaining the high quality of the original image.

[0028] The collected video images are processed in real time by the in-vehicle computing platform, mainly including steps such as denoising and enhancement. The system performs denoising on the video images to reduce the impact of environmental noise on image quality, such as image noise under vehicle vibration, sensor interference, or low-light conditions; then, the contrast of the image is enhanced through histogram equalization technology, making road boundaries, lane lines, and markers clearer; in addition, to further improve the usability of the image, the system also performs local brightness adjustment on the areas with abnormal illumination in the video images to ensure that the key features of the road scene can be completely retained. Based on the video image processing, the system extracts key frames from the video stream as the final road scene images. The extraction process of key frames is achieved through an inter-frame change detection algorithm, that is, analyzing the changes in scene features in consecutive video frames and extracting the frames containing significant environmental changes or key road information.

[0029] The target in the front area is implemented using a target detection network, and the target detection network includes a backbone network, a feature enhancement module, a multi-scale feature fusion module, and a target detection head;

[0030] Among them, the backbone network obtains four different-depth feature maps of the road scene image through four sequentially connected feature extraction modules The feature extraction module is defined as:

[0031]

[0032] Among them, Conv is the convolution calculation, the subscript is the convolution kernel size, and the superscript is the number of consecutive executions. are the intermediate feature maps of the feature extraction module respectively. and are the input and output of the feature extraction module respectively. is the per-pixel multiplication.

[0033] In the temporary maintenance scenario of unstructured roads, the detection of targets such as fences, roadblocks, construction workers, and vehicles is relatively complex. The reason is that there are significant differences in the shapes, sizes, texture features, and distribution patterns of these targets, and the scene background may also be relatively chaotic. To solve this problem, the present invention adopts a parallel feature extraction structure to fully explore the features of each target. First, through simple convolution operations, the present invention can quickly extract the basic edge features and low-level semantic information of the target. This simple structure can efficiently capture the features of targets with regular shapes or clear boundaries, such as fences and roadblocks, and lay a foundation for subsequent multi-scale feature extraction.

[0034] At the same time, to solve the problem of different scales and sizes of targets such as fences, roadblocks, construction workers, and vehicles, the present invention also designs a multi-scale extraction method to obtain The single-convolution path can extract the local features of the target respectively, which is suitable for capturing the detailed features of small targets such as construction workers and distant vehicles; the double-convolution path enhances the perception ability of medium-scale targets by using a deeper receptive field, such as fences or medium-sized roadblocks; the triple-convolution path extracts deeper semantic features, which helps to better understand complex targets, such as large roadblocks or overlapping multi-targets in the background. Multi-size target extraction can effectively cope with the problem of large-scale changes of targets in unstructured roads, and at the same time, it can retain enough detailed information.

[0035] Furthermore, the fences in the temporary maintenance scenario usually show a long strip distribution, and the distribution of vehicles also has a certain directionality. To improve the direction sensitivity and spatial perception ability of feature extraction, the present invention also extracts feature maps as supplements. Among them, the 1*3 and 3*1 directional convolution calculations can flexibly extract the horizontal, vertical, and oblique features of the target, thereby enhancing the recognition ability of different target forms.

[0036] The feature maps are respectively input into the feature enhancement module for foreground feature enhancement to obtain

[0037] Among them, are the input and output of the feature enhancement module respectively, w en is the weight feature map, is the per-pixel multiplication, and BN, σ, and FC are normalization, activation function, and fully connected respectively;

[0038] In the scenario of a temporary maintenance road in an unstructured section, background information such as the sky and asphalt road surface usually occupies most of the image area. However, this background information is irrelevant to the object detection task and may even interfere with the accurate recognition of foreground objects. Therefore, the present invention generates a weight map for the extracted foreground features, guides the attention of the foreground and background to suppress the background information. Through this method, the interference of the background area on feature extraction can be effectively reduced, thereby enhancing the feature enhancement ability of the subsequent network for foreground object features. In addition, this attention guidance mechanism can not only concentrate the computing resources of the network on the key areas where the foreground objects are located, but also help the network extract valuable context clues from the background information. For example, in the temporary maintenance scenario, the relative position relationship between the asphalt road surface and the foreground objects is usually relatively fixed. For example, the relationship between roadblocks or fences and the road surface. By suppressing the irrelevant road surface texture information while retaining its associated features with the foreground objects, it can better assist the object detection task.

[0039] The multi-scale feature fusion module fuses the feature maps and to obtain the feature map and then fuses it with to obtain the fused output feature map

[0040]

[0041] Among them, are the two input feature maps of the fusion module respectively. When the feature maps and are fused When the feature maps and are fused,

[0042] The object detection head is based on to obtain the detection result map F containing the object category and the corresponding position information det ; the detected set of K objects is Each object o k includes: the center coordinates of the bounding box (x k , y k ), and the object category c k ∈ {fence, construction worker, moving vehicle, roadblock}.

[0043] Divide the passable area based on the road scene image;

[0044] The present invention uses the detection result map to guide the division of the passable area of the road production and economic image. First, a composite weight map is constructed for each pixel (x, y):

[0045]

[0046] Wherein: is the target distance, σ b = 100 is the global background Gaussian radius; σ c is the class radius parameter;

[0047] In an implementable exemplary embodiment, the Gaussian parameters related to the class can be specifically defined and σ c : suppression amplitude α c : enclosure = -1.2, construction worker = 0.8, roadblock = -1.0, promotion amplitude β c : moving vehicle = 0.6, action radius σ c : enclosure = 20px, construction worker = 30px, moving vehicle = 10px, roadblock = 40px; An encoder is used to extract features from the road scene image to obtain three-level feature encoding maps Downsample the composite weight map to guide the feature map to perform segmentation on key areas;

[0048]

[0049] Wherein, DownSample downsamples the composite weight map W to be consistent with the size, and are the (m - 1)-th and n-th attention-guided feature maps respectively; μ and σ are used to calculate the channel mean and standard deviation respectively; Input it into a decoding network symmetric to the encoder structure to obtain the semantic segmentation map of the passable area in the road scene image;

[0050] During the process of dividing the passable area, a Gaussian weight map is generated by using the extracted target to guide the semantic segmentation of the passable area. In the temporary maintenance scenario of unstructured sections, the boundary of the passable area is usually interfered by complex backgrounds and dynamic targets. By generating a Gaussian weight map, the attention of the network can be effectively focused on the key areas associated with the target, thereby enhancing the perception ability of the segmentation for the passable area and filtering out the interference of the background area or irrelevant information on the segmentation result. The generation of the Gaussian weight map is based on the target information extracted previously, and can dynamically adjust the weight distribution according to the spatial distribution of the target. Specifically, higher weights are assigned to the areas close to construction workers, enclosures or roadblocks, while lower weights are assigned to the background areas far from these targets, thereby guiding the semantic segmentation to pay more attention to the potential passable areas.

[0051] In addition, in the scenario of temporary repair, the shape of the passable area may change due to the complexity of the road conditions or the diversity of the obstacle distribution. For example, some areas may be occupied by roadblocks or fences, making it difficult for traditional segmentation methods to accurately divide the passing path. Guided by the Gaussian weight map, the segmentation process can automatically adapt to these complex changes, dynamically adjust the degree of attention to different areas, and thus more accurately identify the actual passable road area. Especially when construction workers or vehicles are close to the boundary of the passable area, the Gaussian weight map can effectively highlight the saliency of these areas and help the network more accurately segment the safe passing path.

[0052] Construct virtual lane lines based on the passable area and the detected vehicle position information;

[0053] The set of vehicles detected by the object detection network is where N is the number of vehicles, y i and w i , h i respectively represent the coordinates, target width, and height of the v i -th vehicle;

[0054] Based on the vehicle set query the set of unoccluded vehicles and generate an initial lane for each vehicle in :

[0055]

[0056] where σ x is the lateral position sensitivity coefficient, τ o is the occlusion determination threshold, represents the t-th vehicle in the set , T is the number of vehicles in the set , L t is the t-th lane;

[0057] For any vehicle to be assigned assigned to lane L t should satisfy the condition:

[0058]

[0059] where W lane = 3.75, W lane is the standard lane width, α is the perspective effect compensation coefficient. For each lane L t , sort by the longitudinal coordinate:

[0060] where represents the p-th vehicle belonging to the t-th lane ordinate;

[0061] Based on the vehicle position on lane L t generate the t-th virtual lane line:

[0062] β is the lane line curvature compensation coefficient, Y is the maximum visible distance, and y is the distance extended from the current position to the forward road;

[0063] In the construction area, due to the setting of fences, roadblocks or other temporary facilities, the original lane lines may be partially blocked or completely disappear. Even in some cases, the lane lines may no longer be able to reflect the real lane conditions. If the autonomous driving system still relies on these physical lane lines for path planning, it may lead to wrong judgments, causing the vehicle to drive into a dangerous area or deviate from the safe driving path unnecessarily. However, by analyzing the occlusion relationship between vehicles, the actual lane structure can be inferred from the vehicle distribution. For example, if there is occlusion between two vehicles and their lateral position deviation is very small, it can be judged with a high confidence that they are in the same lane. This judgment based on the position relationship of the targets can effectively replace the invalid static lane line information.

[0064] Generating virtual lane lines based on the position information of vehicles on the same lane also improves the adaptability and robustness of the system. The virtual lane lines no longer rely on the physical lane lines in the construction area, but reflect the real driving path of the road by analyzing the actual vehicle distribution. For example, in the construction area, the original lane may be occupied by fences, and the vehicle needs to drive along the edge of the construction area or other temporarily open channels. At this time, the virtual lane lines generated by analyzing the position information of vehicles on the same lane can accurately describe the current driving trajectory and avoid path planning errors caused by the misleading of the original lane lines.

[0065] Construct an autonomous driving decision according to the virtual lane line in combination with the passable area and the area target.

[0066] Based on the area target recognition result, circular high-risk areas with a radius of 50 cm are generated at the positions of the fences and roadblocks, and their risk values decay exponentially with the distance from the center point; a dynamic elliptical risk field is generated in the construction worker area, and the long axis direction is perpendicular to the driving direction of the vehicle;

[0067] For the high-risk target area, translate the original virtual lane line to the opposite side to form an avoidance area, and at the same time compress the width of the adjacent lane to ensure a safe distance;

[0068] For the construction worker area, a dynamic right-of-way switching point is set 3 meters in front of the position of the construction worker target, triggering a temporary bifurcation of the lane line, interrupting the current lane and activating a detour guiding line.

[0069] For virtual lane lines, map the virtual lane line coordinates to the passable semantic segmentation map. If the overlap rate with the non-passable area exceeds 15%, start path reconstruction. During path reconstruction, attempt to offset the current lane line parallelly, with the offset step increasing by 0.2 meters, and the maximum offset not exceeding the standard lane width. If the conditions are still not met, mark this section as a restricted passing area and trigger the downgraded control mode.

[0070] The present invention provides a multi-sensor fusion collaborative autonomous driving decision-making system for new energy vehicles facing complex road conditions. The system includes:

[0071] An image acquisition device: Collect road scene images through a camera, identify the targets in the front area, and determine entry into a temporary construction area.

[0072] A semantic segmentation module: The semantic segmentation module divides the passable area based on the road scene image. A virtual lane line generation module: The virtual lane line generation module constructs virtual lane lines based on the passable area and the detected vehicle position information.

[0073] An autonomous driving decision-making module: The autonomous driving decision-making module constructs an autonomous driving decision based on the virtual lane line in combination with the passable area and the area targets.

[0074] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-mentioned multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions.

[0075] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions.

[0076] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0077] In this specification, the same or similar parts among the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments described later, the description is relatively simple, and the relevant parts can be referred to the partial description of the foregoing embodiments.

[0078] As described above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions, characterized by: The method comprises: S1: The camera collects road scene images and identifies targets in the front area to determine whether to enter the temporary construction area; S2: Divide the drivable area based on the road scene image; The detection result map is used to guide the road scene image to divide the passable area. The detection result map is obtained by the target detection head based on the output feature map of the multi-scale feature fusion module to obtain the detection result map containing the target category and the corresponding position information. The set of K targets detected by the target detection head is Each target k Contains the coordinates of the center of the bounding box (x k ,y k ) and target category c k ∈{enclosure, construction workers, moving vehicles, roadblocks}; Construct a composite weight map for each pixel (x, y): is the target distance, σ b is the global background Gaussian radius; σ c is the category radius parameter; is the category-dependent Gaussian parameter; The encoder is used to extract features from the road scene image and obtain three levels of feature coding maps: Downsample the composite weight map to guide the feature map to segment key areas; DownSample is to downsample the composite weight map W to The sizes are consistent, μ and σ are the mean and standard deviation of the calculation channels respectively; Conv is the convolution calculation, and the subscript is the convolution kernel size. The semantic segmentation map of the drivable area in the road scene image is obtained by inputting it into the decoding network which is symmetrical with the encoder structure; S3: construct virtual lane lines based on the passable area and the detected vehicle position information; The constructing of the virtual lane includes generating an initialization lane including: The set of vehicles detected by the target detection network is N is the number of vehicles, x i ,y i and w i ,h i Respectively represent the vth i The coordinates of each vehicle and the target width and height; based on the vehicle set Query the unobstructed vehicle set And Generate an initial lane for each vehicle: σ x is the lateral position sensitivity coefficient, τ o is the occlusion determination threshold, Representing a collection The t-th vehicle in the set, T is Number of vehicles in L t is the tth lane; S4: Build autonomous driving decisions based on virtual lane lines combined with drivable areas and regional targets.

2. The multi-sensor fusion collaborative automatic driving decision-making method for complex road conditions according to claim 1 is characterized in that: The targets in the front area are realized by using a target detection network, which includes a backbone network, a feature enhancement module, a multi-scale feature fusion module and a target detection head. The targets in the front area include: fences, construction workers, moving vehicles and roadblocks.

3. The multi-sensor fusion collaborative automatic driving decision-making method for complex road conditions according to claim 2, characterized in that: The backbone network consists of four sequentially connected feature extraction modules to obtain four different depth feature maps of the road scene image. The feature extraction module is defined as: Conv is the convolution calculation, the subscript is the convolution kernel size, and the superscript is the number of consecutive executions. are the intermediate feature maps of the feature extraction module, and are the input and output of the feature extraction module respectively. It is pixel-by-pixel addition.

4. The multi-sensor fusion collaborative automatic driving decision-making method for complex road conditions according to claim 3, characterized in that: The feature map They are input into the feature enhancement module to enhance the foreground features, and the are the input and output of the feature enhancement module, respectively, and w en is the weight feature map, is pixel-by-pixel multiplication, BN, σ, and FC are normalization, activation function, and full connection, respectively.

5. The multi-sensor fusion collaborative automatic driving decision-making method for complex road conditions according to claim 4 is characterized in that: Constructing a virtual lane also includes: Assigned to Lane L t The following conditions should be met: Among them, W lane is the standard lane width, α is the perspective effect compensation coefficient, For each lane L t , sorted by vertical coordinate: in represents the pth vehicle belonging to the tth lane The vertical coordinate of Based on Lane L t The vehicle position is used to generate the tth virtual lane line: β is the lane curvature compensation coefficient, Y is the maximum visible distance, and y is the distance from the current position to the road ahead.

6. A multi-sensor fusion collaborative automatic driving decision system for complex road conditions, used to execute the multi-sensor fusion collaborative automatic driving decision method for complex road conditions as claimed in any one of claims 1 to 5, characterized in that The system includes: Image acquisition equipment: collects road scene images through cameras and identifies targets in the front area to determine whether to enter the temporary construction area; Semantic segmentation module: the semantic segmentation module divides the drivable area based on the road scene image; Virtual lane line generation module: The virtual lane line generation module constructs virtual lane lines based on the passable area and the detected vehicle position information; Autonomous driving decision module: The autonomous driving decision module builds autonomous driving decisions based on virtual lane lines combined with drivable areas and regional targets.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method for multi-sensor fusion collaborative autonomous driving decision-making for complex road conditions as described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the multi-sensor fusion collaborative autonomous driving decision-making method for complex road conditions as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • System and method for automatically driving to pass through construction road section

    CN114506347A

  • Lane detection method and system based on vision and lidar multi-level fusion

    US10929694B1