Map construction method and system integrating traffic sign recognition, lane detection and rule integration

By integrating traffic sign recognition and lane detection, and combining visual and text information to build a dynamic rule-based map method, the insufficient online construction of high-precision maps at the traffic rule layer is solved, and the safety and decision-making reliability of autonomous driving are improved.

CN120747918APending Publication Date: 2025-10-03BEIHANG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510899971.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing high-precision maps have weak online construction capabilities at the traffic rules layer, and traffic sign detection and lane line recognition are not robust in severe weather and occlusion scenarios, making it difficult to accurately associate rules with lanes, affecting the integrity of the autonomous driving decision-making logic.

Method used

By acquiring data from the vehicle's camera and lidar equipment, the system uses preset detection models to identify traffic signs and lane division lines, combines visual features and text information to build dynamic rules, and generates online maps based on real-time positioning data.

Benefits of technology

It improves the reliability and real-time performance of map construction, prevents autonomous vehicles from mistakenly entering restricted areas or violating traffic regulations, and enhances safety and decision-making reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747918A_ABST
    Figure CN120747918A_ABST
Patent Text Reader

Abstract

The invention provides a map construction method and system fusing traffic sign recognition, lane detection and rule integration, and the method comprises the steps: obtaining a detection result of each traffic sign in a collected image through a preset detection model, obtaining a segmentation result of a lane line in the image, and obtaining the lane information through the combination of the segmentation result, point cloud and text prompt. And obtaining visual features and text information of the traffic signs based on the detection results of the traffic signs, obtaining rule information by combining the visual features and the text information, and obtaining dynamic rules corresponding to the lane lines based on the lane information and the rule information. And constructing an online map by combining the lane information, the dynamic rule and the real-time positioning data of the vehicle. According to the scheme, by introducing rule extraction of the traffic signs and correlation reasoning of the lane lines, reliability and real-time performance of map construction are improved, the situation that the automatic driving vehicle enters the forbidden area by mistake or violates traffic rules is avoided, and safety and decision reliability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of autonomous driving perception technology, and more specifically, to a map construction method and system that integrates traffic sign recognition, lane detection, and rule integration. Background Art

[0002] The construction and real-time updating of high-precision maps are critical components of autonomous driving technology. Currently, existing HD maps mostly focus on the construction of the geometry layer (e.g., lane markings and curbs) and the connectivity layer (e.g., lane topology). However, the online construction of the traffic rules layer (e.g., speed limits, dedicated lanes, etc.) still relies on offline data, resulting in weak real-time update capabilities.

[0003] Furthermore, traffic sign detection and lane recognition face significant robustness challenges in adverse weather and occlusion scenarios, with traditional perception models prone to missed or false detections. Furthermore, existing methods typically describe traffic signs through classification labels, but lack structured rule representation and struggle to accurately associate rules with specific lanes, compromising the integrity of autonomous driving decision-making logic. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a map construction method and system that integrates traffic sign recognition, lane detection, and rule integration to improve the reliability and real-time performance of map construction, avoid autonomous vehicles from mistakenly entering restricted areas or violating traffic rules, and improve safety and decision-making reliability.

[0005] In a first aspect, the present invention provides a map construction method integrating traffic sign recognition, lane detection, and rule integration, the method comprising: Obtain images captured by the vehicle's camera and point clouds collected by the lidar; Preprocessing the image and the point cloud; Obtaining detection results of each traffic sign in the image using a preset detection model; Detecting and obtaining a segmentation result of lane lines in the image, and combining the segmentation result, the point cloud, and text prompts to obtain lane information; Obtaining visual features and text information of each traffic sign based on a detection result of each traffic sign in the image, obtaining rule information by combining the visual features and the text information, and obtaining a dynamic rule corresponding to each lane line based on the lane information and the rule information; An online map is constructed by combining the lane information, dynamic rules and real-time positioning data of the vehicle.

[0006] In an optional embodiment, the step of preprocessing the image and the point cloud includes: Performing time alignment processing on the collected images and point clouds based on the timestamps of the camera device and the lidar; performing radial distortion correction and tangential distortion correction on the image; The image is divided into a plurality of sub-images, and contrast adjustment and noise suppression processing are performed on each of the sub-images.

[0007] In an optional embodiment, the step of performing time alignment processing on the collected images and point clouds based on the timestamps of the camera device and the lidar includes: Based on the timestamps of the camera device and the laser radar, the optimal reference time is determined by the least squares method to minimize the sum of squares of the deviations between the timestamp of the camera device and the reference time and the timestamp of the laser radar and the reference time.

[0008] In an optional embodiment, the detection model includes a first detection model, a second detection model and a convolutional neural network model; The detection model is trained in the following way: Obtaining a sample image, wherein the sample image has a traffic sign annotated label; Importing the sample image into a first detection model, and outputting prediction information of traffic signs in the sample image; executing training of the first detection model based on a first loss function constructed based on the prediction information and the annotated labels, where the first loss function is constructed by an intersection-over-union ratio of a predicted box in the prediction information and a true box in the annotated labels, and a Euclidean distance between center points; When the confidence level of the prediction information is less than a preset threshold, the second detection model and the convolutional neural network model are trained again based on the cropped image of the traffic sign with a confidence level less than the preset threshold until the preset requirements are met.

[0009] In an optional embodiment, the step of further training the second detection model and the convolutional neural network model based on the cropped image of the traffic sign having a confidence level less than a preset threshold until preset requirements are met includes: Importing the cropped image of the traffic sign with a confidence score less than a preset threshold into the second detection model, and training the second detection model based on a second loss function constructed based on the classification loss and the calibration box loss; The sign area corresponding to the traffic sign output by the second detection model is imported into the convolutional neural network model, and the convolutional neural network model is trained based on the predicted category of the traffic sign output and the labeled category in the labeled label until the preset requirements are met.

[0010] In an optional embodiment, the step of combining the segmentation result, the point cloud, and the text prompt to obtain lane information includes: Projecting the point cloud into a feature space to obtain point cloud features, and mapping text prompts into semantic vectors; extracting deep features from the image based on the lane line segmentation result; Combining the point cloud features, the semantic vectors, and the deep features to generate fusion features; Completion information is generated based on the fusion feature, and the completion information is fused with the segmentation result to obtain completed lane information.

[0011] In an optional embodiment, the step of extracting deep features from the image based on the lane line segmentation result includes: Outputting binary mask information of the lane line in the image based on the segmentation result of the lane line in the image; Performing curve fitting on the lane line based on the binary mask information to obtain geometric information of the lane line; Deep features are extracted based on the geometric information of the lane lines.

[0012] In an optional embodiment, the step of obtaining a dynamic rule corresponding to each lane line based on the lane information and the rule information includes: Obtaining geometric features and topological features based on the lane information encoding; Calculating the similarity between the geometric features and the rule information, and the similarity between the topological features and the rule information; Based on the similarity between the geometric features and the rule information, and the similarity between the topological features and the rule information, a dynamic rule corresponding to each lane line is constructed.

[0013] In an optional embodiment, the step of constructing an online map by combining the lane information, dynamic rules, and real-time positioning data of the vehicle includes: Building a local map by combining the lane information, dynamic rules, and real-time positioning data of the vehicle; A logical conflict detection is performed on the local map based on information in a preset original data map, and the local map is corrected based on the detection result to obtain a corrected online map.

[0014] In a second aspect, the present invention provides a map construction system integrating traffic sign recognition, lane detection, and rule integration, the system comprising: An acquisition module is used to acquire images captured by the camera equipment on the vehicle and point clouds collected by the lidar; A preprocessing module, configured to preprocess the image and the point cloud; a first detection module, configured to obtain a detection result of each traffic sign in the image using a preset detection model; A second detection module is configured to detect and obtain a segmentation result of lane lines in the image, and obtain lane information by combining the segmentation result, the point cloud, and text prompts; a combining module, configured to obtain visual features and text information of each traffic sign based on a detection result of each traffic sign in the image, obtain rule information by combining the visual features and the text information, and obtain a dynamic rule corresponding to each lane line based on the lane information and the rule information; A construction module is used to construct an online map by combining the lane information, dynamic rules and real-time positioning data of the vehicle.

[0015] The present invention provides a map construction method and system that integrates traffic sign recognition, lane detection, and rule integration. By acquiring images captured by a vehicle's camera and a point cloud collected by a laser radar (LiDAR), a preset detection model is used to obtain detection results for each traffic sign in the image. Lane segmentation results are then detected and obtained, and lane information is then combined with the segmentation results, point cloud, and textual prompts. Based on the detection results for each traffic sign, visual features and textual information are obtained for the traffic sign. Rule information is then obtained based on the visual features and textual information. Dynamic rules corresponding to each lane line are then obtained based on the lane information and rule information. An online map is then constructed by combining the lane information, dynamic rules, and real-time vehicle positioning data.

[0016] In this solution, by introducing rule extraction of traffic signs and associative reasoning with lane lines, the reliability and real-time performance of map construction are improved, preventing autonomous vehicles from mistakenly entering restricted areas or violating traffic regulations, thereby improving safety and decision-making reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A flowchart of a map construction method integrating traffic sign recognition, lane detection, and rule integration provided by an embodiment of the present invention; Figure 2 for Figure 1 Flowchart of the sub-steps included in S12; Figure 3A flowchart of a model training method in a map construction method provided in an embodiment of the present invention; Figure 4 for Figure 1 Flowchart of the sub-steps included in S14; Figure 5 for Figure 1 Flowchart of the sub-steps included in S15; Figure 6 for Figure 1 A flowchart of the sub-steps included in S16; Figure 7 A functional module block diagram of a map building system integrating traffic sign recognition, lane detection, and rule integration provided by an embodiment of the present invention; Figure 8 This is a structural block diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.

[0020] See also Figure 1 This is a flow chart of a map construction method integrating traffic sign recognition, lane detection, and rule integration, provided in an embodiment of the present invention. This map construction method can be performed by a map construction system integrating traffic sign recognition, lane detection, and rule integration. This map construction system can be implemented in software and / or hardware and can be configured in an electronic device, such as a computer device or server, for example, a server in a backend control platform. The detailed steps of this map construction method are described below.

[0021] S11, obtaining images captured by the camera equipment on the vehicle and point clouds captured by the laser radar.

[0022] S12: Preprocess the image and the point cloud.

[0023] S13, using a preset detection model to obtain detection results of each traffic sign in the image.

[0024] S14, detecting and obtaining a segmentation result of the lane lines in the image, and combining the segmentation result, the point cloud, and the text prompt to obtain lane information.

[0025] S15, based on the detection results of each traffic sign in the image, obtain visual features and text information of each traffic sign, combine the visual features and the text information to obtain rule information, and obtain dynamic rules corresponding to each lane line based on the lane information and the rule information.

[0026] S16, constructing an online map by combining the lane information, dynamic rules and real-time positioning data of the vehicle.

[0027] The map construction method provided in this embodiment integrates traffic sign recognition, lane detection, and rule integration. By introducing rule extraction of traffic signs and reasoning about their association with lane lines, it improves the reliability and real-time performance of map construction, prevents autonomous vehicles from mistakenly entering restricted areas or violating traffic rules, and improves safety and decision-making reliability.

[0028] The specific implementation of each of the above steps in this embodiment is described in detail below.

[0029] In this embodiment, the vehicle is equipped with multiple sensors, including cameras and lidar, which can collect environmental data around the vehicle, including images and point clouds. The collected images and point clouds need to be preprocessed to facilitate subsequent analysis and processing.

[0030] In this embodiment, please refer to Figure 2 , the steps of preprocessing images and point clouds can be achieved in the following ways: S121, performing time alignment processing on the collected images and point clouds based on the timestamps of the camera device and the lidar.

[0031] S122: Perform radial distortion correction and tangential distortion correction on the image.

[0032] S123: Divide the image into multiple sub-images, and perform contrast adjustment and noise suppression processing on each sub-image.

[0033] To ensure the temporal consistency of images and point clouds from different devices, such as cameras and lidars, we first need to perform temporal alignment on the images and point clouds. Specifically, this can be achieved by: Based on the timestamps of the camera device and the laser radar, the optimal reference time is determined by the least squares method to minimize the sum of squares of the deviations between the timestamp of the camera device and the reference time and the timestamp of the laser radar and the reference time.

[0034] In this embodiment, hardware timestamps are used to align time series data from cameras, lidars, and other sensors, keeping the time error within 10ms. Specifically, the least squares method is used to find an optimal reference time that minimizes the sum of the squared deviations between the timestamps of all sensors (including cameras and lidars) and this reference time:

[0035] in, is the unified reference time after alignment, The timestamp of each sensor (including camera and lidar).

[0036] On this basis, the image is corrected for radial and tangential distortion based on the camera calibration parameters and a distortion correction model, such as the Brown-Conrady model. The formula is as follows:

[0037] in, is the normalized coordinate after correction (only the x direction is taken as an example, the y direction is the same), is the radial distortion coefficient, is the tangential distortion coefficient, .

[0038] The CLAHE algorithm is then used to partition the image into multiple sub-images. Each sub-image is then independently contrast-adjusted and noise suppression is performed, enhancing the visibility of dark area details while avoiding overexposure. To further enhance generalization in complex environments, Gaussian noise and motion blur can be dynamically injected during the training phase of the distortion correction model to simulate scenes such as rain, fog, and shaking.

[0039] After preprocessing the image and point cloud, the detection results of each traffic sign in the image can be obtained using a preset detection model, where the detection model is pre-trained based on a large number of annotated images.

[0040] The detection model includes the first detection model, the second detection model and the convolutional neural network model. Figure 3 , the following describes how to implement pre-training of the detection model: S21, obtaining a sample image, wherein the sample image has a traffic sign label.

[0041] S22: Import the sample image into a first detection model, and output prediction information of the traffic sign in the sample image.

[0042] S23, performing training of the first detection model based on a first loss function constructed based on the prediction information and the annotated label, where the first loss function is constructed by the intersection-over-union ratio of the predicted box in the prediction information and the true box in the annotated label, and the Euclidean distance between the center points.

[0043] S24, when the confidence level of the prediction information is less than a preset threshold, the second detection model and the convolutional neural network model are trained again based on the cropped image of the traffic sign with a confidence level less than the preset threshold until the preset requirements are met.

[0044] In this embodiment, each sample image is first labeled in the following manner: Mark the center coordinates of the bounding box of the traffic sign in the sample image , width and height And category c. Lane lines are represented by polygon vertex sequences The traffic sign is represented by a smooth Bezier curve and associated with semantic attributes. The optical character recognition (OCR) method is used to extract the rule text represented by the traffic sign and parse it into structured rules, outputting a string. .

[0045] Among them, rule text refers to the text information directly expressed on traffic signs with clear traffic restriction meanings, which can specifically include the following: (1) Direct restriction text: such as speed limit signs, lane function instructions, prohibition signs, etc. (2) Compound rule text: such as spatiotemporal restriction rules (similar to "7:00-9:00 bus only"), conditional rules, etc. (3) Special scenario text: such as temporary construction signs, large-scale event signs, etc.

[0046] Multi-dimensional annotation of sample images in the above manner can provide rich supervisory signals for subsequent detection, classification, and rule understanding tasks.

[0047] The detection model is trained using sample images annotated in the above manner. The detection model includes three levels: a first detection model, a second detection model, and a convolutional neural network model. The first detection model can be a YOLOv8 model, which can be used for real-time traffic sign detection. The second detection model can be an RT-DETR model, which can be used to enhance detection robustness for occluded images in complex scenes. The convolutional neural network model can use a ResNet-50 network, which can be used to output high-confidence detection boxes and sign categories.

[0048] Specifically, the preprocessed and annotated sample images can be input into the YOLOv8 model for preliminary detection, and a set of candidate boxes (i.e., predicted information of traffic signs) can be output through regression analysis:

[0049] in, is the category confidence. The prediction information includes the predicted box, and the annotation label includes the true box. Based on the predicted box and the true box, CLoU Loss is used to construct the first loss function for optimizing bounding box regression. The first loss function is as follows:

[0050] Among them, IoU represents the intersection-over-union ratio between the predicted box and the real box. The center point of the prediction box bThe center point of the real frame The Euclidean distance between d is the minimum bounding box diagonal length, α is the weight coefficient used to balance the effect of the aspect ratio term, v is a measure of aspect ratio consistency.

[0051] For low confidence detection results in complex scenarios (such as <0.5), the system automatically triggers the RT-DETR secondary detection mechanism. The model receives the cropped image of the traffic label output by YOLOv8 with a confidence score less than the preset threshold. , then train the subsequent second detection model and convolutional neural network model based on the cropped image. Specifically, this step can be achieved by: The cropped image of the traffic sign with a confidence level less than a preset threshold is imported into the second detection model, and the second detection model is trained based on a second loss function constructed based on the classification loss and the calibration box loss; the sign area corresponding to the traffic sign output by the second detection model is imported into the convolutional neural network model, and the convolutional neural network model is trained based on the output predicted category of the traffic sign and the labeled category in the labeled label until the preset requirements are met.

[0052] In this embodiment, during the training of the second detection model, the global attention mechanism based on Transformer uses the Hungarian algorithm to match the predicted box with the true box, and the constructed second loss function is as follows:

[0053] in, is the weight coefficient of the classification loss and the calibration box loss, which is used to balance the two contributions. The classification loss and the calibration box loss can be implemented using the losses commonly used in the art.

[0054] By analyzing the contextual relationships between different image regions, the robustness of detection of occluded and blurred signs can be effectively improved.

[0055] Finally, during the training of the convolutional neural network model, the traffic sign area detected by the second detection model is The data is fed into a convolutional neural network model, such as the ResNet-50 network, for fine classification. Deep feature extraction is used to output the predicted category and corresponding probability of the traffic sign, forming the category probability ŷ:

[0056] The convolutional neural network model is trained by combining the labeled categories and predicted categories in the annotation labels until the preset requirements are met.

[0057] In this embodiment, the three parts of the detection model adopt a dynamic collaboration mechanism. Among them, the high-confidence results of YOLOv8 can be directly output in real time, and the subsequent verification process is only initiated for uncertain samples. This intelligent diversion strategy enables the system to maintain high accuracy while reducing the computational load.

[0058] The pre-trained detection model can be applied to real-time detection. Specifically, the pre-trained detection model is used to detect traffic signs in real-time images, generating detection results. The detection results include the category of each traffic sign and its calibrated bounding box. The traffic sign category and its calibrated bounding box information are then combined to provide structured traffic sign data (coordinates, category, and confidence level) for subsequent map construction, ensuring accurate generation of the traffic rules layer.

[0059] In addition to detecting traffic signs, it is also necessary to segment the lane lines in the image to obtain segmentation results. The lane information is obtained by combining the segmentation results, point cloud, and text prompts. For details, please refer to Figure 4 , this step can be achieved by: S141 , projecting the point cloud into a feature space to obtain point cloud features, and mapping the text prompt into a semantic vector.

[0060] S142: Extract deep features from the image based on the lane line segmentation result.

[0061] S143: Generate a fusion feature by combining the point cloud feature, the semantic vector, and the deep feature.

[0062] S144 , generating completion information based on the fusion feature, and fusing the completion information with the segmentation result to obtain completed lane information.

[0063] In this embodiment, the lane segmentation results, point cloud BEV, and text prompts are first input into a large language model, such as the LLAMA-2 model, for cross-modal feature alignment. The text prompts are the regular text of the traffic signs obtained above.

[0064] In this embodiment, the EVA model can be used to project the BEV point cloud into the feature space In addition, the LLAMA-2-7B model can be used to Mapping to semantic vectors , to capture the logical relationships in language instructions.

[0065] At the same time, the deep features of the image are extracted based on the segmentation results of the lane lines through the CNN model , to achieve alignment of geometric and surface features. In this embodiment, the step of extracting deep features from the image based on the lane line segmentation result can be achieved by: Based on the segmentation result of the lane line in the image, binary mask information of the lane line in the image is output; based on the binary mask information, curve fitting is performed on the lane line to obtain geometric information of the lane line; and based on the geometric information of the lane line, deep features are extracted to obtain.

[0066] In this embodiment, the pre-processed image is input into the CNN network, and pixel-level semantic segmentation is performed through the encoder-decoder structure in the CNN network to output the binary mask of the lane line. , where 1 represents a lane line pixel. The segmentation result C not only contains the geometric information of the lane line, but also distinguishes the lane type by different pixel values. Based on the binary mask information of the lane line, a polynomial curve is fitted to the lane line to parameterize the lane line geometry:

[0067] in, x Represents the coordinates of the pixel points on the lane line, y Represents the curve information after fitting, a 0. a 1. a 2. a 3 represents the fitting parameter. The deep features of the lane line are obtained based on the geometric information of the lane line. .

[0068] On the basis of the above, the attention mechanism is used to dynamically weight multimodal features, including point cloud features, semantic vectors, and deep features, to generate fusion features:

[0069] in, , represents comprehensive visual features, , representing semantic features. The LLAMA-2 decoder is based on fusion features Output completion information , parse the completion information into geometric parameters and rasterize it into the completion area and fused with the original segmentation result C:

[0070] Rasterize represents the rasterization process. The final output retains the details detected by visual inspection while completing lane information missing due to occlusion or degradation.

[0071] On the basis of the above, the visual features and text information of each traffic sign are obtained based on the detection results of each traffic sign in the image, and the rule information is obtained by combining the visual features and text information.

[0072] Specifically, this embodiment uses a visual language encoder to deeply analyze the semantic content of traffic signs based on preprocessed images. This semantic content includes a combined understanding of the traffic sign's visual features and the rule text. First, an image classification model, such as the ViT model, can be used to extract the multi-layered visual features of each traffic sign in the image. This is then combined with optical character recognition (OCR) to obtain the textual information.

[0073] Use cross-modal attention mechanism to fuse visual features and text information, and finally output structured rule information , where each rule Represented in the form of key-value pairs. To address the problem of variable-length rule representation, the visual language encoder introduces the [CLS] tag to encode variable-length rule information into a fixed-dimensional feature vector.

[0074] Based on the above, the dynamic rules corresponding to each lane line are obtained based on the lane information and rule information. Figure 5 , this step can be achieved by: S151, obtaining geometric features and topological features based on the lane information encoding.

[0075] S152: Calculate the similarity between the geometric feature and the rule information, and the similarity between the topological feature and the rule information.

[0076] S153: Constructing dynamic rules corresponding to each lane line based on the similarity between the geometric features and the rule information and the similarity between the topological features and the rule information.

[0077] The lane information set L and the rule information set R are input into the map element encoder, and the lane semantics are distinguished through type embedding. Lane semantics include lane functional attributes (such as lane category) and spatiotemporal constraints (such as lane topology and associated traffic rules).

[0078] Based on lane information, the geometric and topological features of the lane are obtained through Transformer encoding. The similarity between the geometric features and the rule information, as well as the similarity between the topological features and the rule information, is calculated. Finally, a bipartite graph is generated that can be directly used for real-time decision-making of the autonomous driving system. ,in Represents the incidence matrix. This bipartite graph is a dynamic association network between rules and lanes, consisting of rule vertices R, lane vertices L (a structured rule set), and an incidence matrix E (a set of lane vectors). The incidence matrix contains m rules and k lanes. An element 1 in the incidence matrix indicates that the rule applies to a lane, while a 0 indicates that the rule has no association with the lane. In other words, this bipartite graph can be used to represent the dynamic rules of each lane segment.

[0079] The difference between the dynamic rules and the aforementioned rule sets is that the rule sets are structured rules extracted directly from a single frame and are valid at the current detection moment. Dynamic rules, on the other hand, are the spatiotemporal fusion of multiple frame rule sets and may be in effect continuously (e.g., requiring a continuous speed limit on a certain road section).

[0080] On this basis, the online map is constructed by combining lane information, dynamic rules and real-time vehicle positioning data. For details, please refer to Figure 6 , this step can be achieved by: S161: Build a local map by combining the lane information, dynamic rules, and real-time positioning data of the vehicle.

[0081] S162: Performing a logic conflict detection on the local map based on information in a preset original data map, and correcting the local map based on the detection result to obtain a corrected online map.

[0082] In this embodiment, by fusing lane information (including lane geometry and topological relationship information) and dynamic rules, using positioning data to stitch together local maps and a multimodal large language model to verify logical consistency, an online high-precision map is generated that includes a geometry layer (lane lines), a connectivity layer (lane topology), and a rule layer (speed limits, dedicated lanes).

[0083] Specifically, lane information (including lane geometry and topology) and dynamic rules are input. Based on the vehicle's real-time positioning data, the detected lanes and rule information are stitched together into a local map, ensuring spatiotemporal consistency of geometry and topology. The local map contains lane lines, lane topology, and the associated dynamic rules.

[0084] Based on the original data map, a large language model is used to analyze multimodal data, detect logical conflicts in local maps, and output a revised online high-precision map:

[0085] Among them, L, T, and R represent the geometric layer, connectivity layer, and regularity layer, respectively.

[0086] The original data map is the input of logical conflict detection, and the final map is output only after LLM verification.

[0087] The conflict detection mechanism is implemented as follows: First, three types of conflicts are identified: (1) conflicts between rules; (2) conflicts between rules and geometry; and (3) conflicts between rules and topology. Then, data from the geometry, topology, and rule layers are integrated with real-time sensor information to generate a conflict score through multimodal feature analysis. Finally, a graded approach is implemented based on the conflict risk level: low-risk rules are automatically corrected, high-risk rules are discarded and trigger manual review, and critical risks trigger vehicle downgrade control.

[0088] The map construction method provided in this embodiment integrates traffic sign recognition, lane detection and rule integration. By introducing traffic sign rule extraction and lane association reasoning, it achieves high-precision and high-efficiency online construction of the traffic rule layer, making up for the shortcomings of existing high-precision maps in real-time updating.

[0089] In addition, by accurately associating structured rules with lanes, autonomous vehicles can be prevented from mistakenly entering restricted areas or violating traffic regulations, effectively improving safety and decision-making reliability.

[0090] Furthermore, by combining visual, textual, and lidar data, the limitations of traditional single sensors are overcome. A multimodal large language model is used for multi-level, deep perception reasoning, enhancing the environmental adaptability of autonomous vehicles and supporting stable perception in extreme conditions such as rain, snow, and low light.

[0091] In summary, the present invention provides a vectorized high-precision map construction solution that integrates traffic sign recognition, lane detection, and rule integration. Through a multimodal reasoning framework, it not only improves the accuracy of environmental perception, but also can adapt to complex traffic environments in real time, and has broad application prospects.

[0092] Based on the same inventive concept, please refer to Figure 7 , an embodiment of the present invention further provides a functional module diagram of a map construction system that integrates traffic sign recognition, lane detection, and rule integration. This embodiment can divide the functional modules of the map construction system according to the above method embodiment. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0093] For example, when each functional module is divided into corresponding functional modules, Figure 7The map construction system shown is only a schematic diagram of the system. The map construction system may include an acquisition module, a preprocessing module, a first detection module, a second detection module, a combination module, and a construction module. The functions of each functional module of the map construction system are described in detail below.

[0094] An acquisition module is used to acquire images captured by the camera equipment on the vehicle and point clouds collected by the lidar; A preprocessing module, configured to preprocess the image and the point cloud; a first detection module, configured to obtain a detection result of each traffic sign in the image using a preset detection model; A second detection module is configured to detect and obtain a segmentation result of lane lines in the image, and obtain lane information by combining the segmentation result, the point cloud, and text prompts; a combining module, configured to obtain visual features and text information of each traffic sign based on a detection result of each traffic sign in the image, obtain rule information by combining the visual features and the text information, and obtain a dynamic rule corresponding to each lane line based on the lane information and the rule information; A construction module is used to construct an online map by combining the lane information, dynamic rules and real-time positioning data of the vehicle.

[0095] The map construction system provided in this embodiment can be used to execute the map construction method under any implementation method of the above embodiments. For details not provided in this embodiment, please refer to the corresponding description of the above embodiments, and this embodiment will not be repeated here.

[0096] See also Figure 8 , is a block diagram of the structure of an electronic device provided in an embodiment of the present invention. This electronic device may be a computer device, server, or the like in an autonomous driving control platform. The electronic device includes a memory, a processor, and a communication module. The memory, processor, and communication module components are electrically connected to each other, directly or indirectly, to enable data transmission or interaction. For example, these components may be electrically connected to each other via one or more communication buses or signal lines.

[0097] Memory is used to store computer programs or data. Memory can include, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM).

[0098] The processor is used to read / write data or programs stored in the memory and execute the map construction method provided by any embodiment of the present invention.

[0099] The communication module is used to establish a communication connection between the electronic device and other communication terminals through the network, and is used to send and receive data through the network.

[0100] It should be understood that Figure 8 The structure shown is only a schematic diagram of the structure of the electronic device. The electronic device may also include Figure 8 More or fewer components than shown, or with Figure 8 Different configurations shown.

[0101] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are executed, the map construction method provided in the above embodiment is implemented.

[0102] Specifically, the computer-readable storage medium can be a general-purpose storage medium, such as a removable disk or hard disk. When the computer program on the computer-readable storage medium is executed, the aforementioned map construction method can be executed. Regarding the processes involved in executing the computer-readable storage medium and its executable instructions, please refer to the relevant descriptions in the aforementioned method embodiments and will not be detailed here.

[0103] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some communication interface, the indirect coupling or communication connection of the device or unit may be electrical, mechanical or other forms.

[0104] In addition, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0105] Furthermore, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0106] It should be noted that if a function is implemented as a software module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0107] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.

[0108] The foregoing description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A map construction method integrating traffic sign recognition, lane detection and rule integration, characterized in that: The method comprises: Obtain images captured by the vehicle's camera and point clouds collected by the lidar; Preprocessing the image and the point cloud; Obtaining detection results of each traffic sign in the image using a preset detection model; Detecting and obtaining a segmentation result of lane lines in the image, and combining the segmentation result, the point cloud, and text prompts to obtain lane information; Obtaining visual features and text information of each traffic sign based on a detection result of each traffic sign in the image, obtaining rule information by combining the visual features and the text information, and obtaining a dynamic rule corresponding to each lane line based on the lane information and the rule information; An online map is constructed by combining the lane information, dynamic rules and real-time positioning data of the vehicle.

2. The map construction method integrating traffic sign recognition, lane detection and rule integration according to claim 1, characterized in that: The step of preprocessing the image and the point cloud comprises: Performing time alignment processing on the collected images and point clouds based on the timestamps of the camera device and the lidar; performing radial distortion correction and tangential distortion correction on the image; The image is divided into a plurality of sub-images, and contrast adjustment and noise suppression processing are performed on each of the sub-images.

3. The map construction method integrating traffic sign recognition, lane detection and rule integration according to claim 2 is characterized in that: The step of performing time alignment processing on the collected images and point clouds based on the timestamps of the camera device and the laser radar comprises: Based on the timestamps of the camera device and the laser radar, the optimal reference time is determined by the least squares method to minimize the sum of squares of the deviations between the timestamp of the camera device and the reference time and the timestamp of the laser radar and the reference time.

4. The map construction method integrating traffic sign recognition, lane detection and rule integration according to claim 1, characterized in that: The detection model includes a first detection model, a second detection model and a convolutional neural network model; The detection model is trained in the following way: Obtaining a sample image, wherein the sample image has a traffic sign annotated label; Importing the sample image into a first detection model, and outputting prediction information of traffic signs in the sample image; executing training of the first detection model based on a first loss function constructed based on the prediction information and the annotated labels, where the first loss function is constructed by an intersection-over-union ratio of a predicted box in the prediction information and a true box in the annotated labels, and a Euclidean distance between center points; When the confidence level of the prediction information is less than a preset threshold, the second detection model and the convolutional neural network model are trained again based on the cropped image of the traffic sign with a confidence level less than the preset threshold until the preset requirements are met.

5. The map construction method integrating traffic sign recognition, lane detection and rule integration according to claim 4 is characterized in that: The step of training the second detection model and the convolutional neural network model based on the cropped image of the traffic sign having a confidence level less than a preset threshold until the preset requirements are met includes: Importing the cropped image of the traffic sign with a confidence score less than a preset threshold into the second detection model, and training the second detection model based on a second loss function constructed based on the classification loss and the calibration box loss; The sign area corresponding to the traffic sign output by the second detection model is imported into the convolutional neural network model, and the convolutional neural network model is trained based on the predicted category of the traffic sign output and the labeled category in the labeled label until the preset requirements are met.

6. The map construction method integrating traffic sign recognition, lane detection and rule integration according to claim 1, characterized in that: The step of combining the segmentation result, the point cloud and the text prompt to obtain lane information includes: Projecting the point cloud into a feature space to obtain point cloud features, and mapping text prompts into semantic vectors; extracting deep features from the image based on the lane line segmentation result; Combining the point cloud features, the semantic vectors, and the deep features to generate fusion features; Completion information is generated based on the fusion feature, and the completion information is fused with the segmentation result to obtain completed lane information.

7. The map construction method integrating traffic sign recognition, lane detection and rule integration according to claim 6 is characterized in that: The step of extracting deep features from the image based on the lane line segmentation result comprises: Outputting binary mask information of the lane line in the image based on the segmentation result of the lane line in the image; Performing curve fitting on the lane line based on the binary mask information to obtain geometric information of the lane line; Deep features are extracted based on the geometric information of the lane lines.

8. The map construction method integrating traffic sign recognition, lane detection and rule integration according to claim 1, characterized in that: The step of obtaining a dynamic rule corresponding to each lane line based on the lane information and the rule information includes: Obtaining geometric features and topological features based on the lane information encoding; Calculating the similarity between the geometric features and the rule information, and the similarity between the topological features and the rule information; Based on the similarity between the geometric features and the rule information, and the similarity between the topological features and the rule information, a dynamic rule corresponding to each lane line is constructed.

9. The map construction method integrating traffic sign recognition, lane detection and rule integration according to claim 1, characterized in that: The step of constructing an online map by combining the lane information, dynamic rules, and real-time positioning data of the vehicle includes: Building a local map by combining the lane information, dynamic rules, and real-time positioning data of the vehicle; A logical conflict detection is performed on the local map based on information in a preset original data map, and the local map is corrected based on the detection result to obtain a corrected online map.

10. A map building system integrating traffic sign recognition, lane detection and rule integration, characterized in that: The system comprises: An acquisition module is used to acquire images captured by the camera equipment on the vehicle and point clouds collected by the lidar; A preprocessing module, configured to preprocess the image and the point cloud; a first detection module, configured to obtain a detection result of each traffic sign in the image using a preset detection model; A second detection module is configured to detect and obtain a segmentation result of lane lines in the image, and obtain lane information by combining the segmentation result, the point cloud, and text prompts; a combining module, configured to obtain visual features and text information of each traffic sign based on a detection result of each traffic sign in the image, obtain rule information by combining the visual features and the text information, and obtain a dynamic rule corresponding to each lane line based on the lane information and the rule information; A construction module is used to construct an online map by combining the lane information, dynamic rules and real-time positioning data of the vehicle.

Citation Information

Patent Citations

  • Automatic driving lane information detection method based on radar point cloud and image fusion

    CN114037969A

  • Multi-target classification method and device based on difficult sample transfer learning

    CN114170532A

  • Lane-level high-precision map construction method and system

    CN116105717A

  • Semantic scene completion method based on image and point cloud fusion in automatic driving scene

    CN116503825A

  • Traffic marker line positioning method, system, equipment and medium

    CN117975396A