Lane attribute creation system, creation method, and computer program product

Through the lane attribute production system, a map element encoder and large language model are used to process multiple modal data, and detailed driving rules and lane correspondence relationships are generated, which solves the problem of lack of detailed rules recognition and matching of traffic rules in the existing technology, and realizes accurate decision-making support for the autonomous driving system.

CN120027805AActive Publication Date: 2025-05-23BEIJING AUTONAVI YUNMAP TECH CO LTD

Patent Information

Application Number
CN202411960752.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-23
Estimated Expiration
2044-12-27

AI Technical Summary

Technical Problem

The high-definition maps in the existing autonomous driving system lack detailed rules identification and matching of traffic signs other than directional signs on the traffic rules layer, and cannot provide the accurate driving rules and lane correspondence required for autonomous driving.

Method used

A lane attribute production system is adopted, which includes a map element encoder and a large language model. By processing vectorized maps, image encoding results and text encoding results, lane attribute data is generated, including the driving rules of the vehicle and the corresponding relationship between driving rules and lanes.

Benefits of technology

The detailed construction of the traffic rules layer is achieved, the accurate driving rules and lane correspondence required by the autonomous driving system is provided, and the decision-making accuracy of autonomous driving and intelligent transportation systems is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120027805A_ABST
    Figure CN120027805A_ABST
Patent Text Reader

Abstract

The invention discloses a lane attribute making system, a lane attribute making method and a computer program product. The lane attribute making system comprises a map element encoder and a large language model. An output layer of the map element encoder is connected to an input layer of the large language model, the map element encoder is used for processing a vectorized map to output a vector encoding result, the vectorized map uses vector features to represent map elements, and the map elements comprise lanes; the input layer of the large language model further receives at least one of the image coding result and the text coding result, and the large language model is configured to generate lane attribute data according to at least one of the image coding result and the text coding result and the vector coding result. According to the scheme provided by the embodiment of the invention, the data of various modes including the vector mode can be processed, so that the driving rules are matched to the corresponding lanes, and accurate and detailed lane attribute data are provided for constructing the traffic rule layer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of map technology. More specifically, the present disclosure relates to a lane attribute making system, a lane attribute making method, and a computer program product. Background Art

[0002] With the vigorous development of technologies such as artificial intelligence and 5G communication technology, autonomous driving and intelligent transportation systems have also progressed to the application stage. The rapid development of autonomous driving and intelligent transportation systems has put forward higher requirements on the reliability and accuracy of navigation data. It requires accurate navigation data to achieve vehicle perception, positioning, path planning and decision control. High-definition (HD) maps, with their detailed representation of road elements, have become an important component to support these systems. The geometry layer of HD maps provides the system with information such as lane dividers and lane centerlines, the connection layer provides the system with lane relationships to facilitate path planning, and the traffic rule layer provides lane-related rule information for the system to make decision control.

[0003] However, the HD maps currently used in autonomous driving systems perform well at the geometry layer and the connection layer, but have the following defects at the traffic rules layer: First, the traffic rules layer it constructs only describes the direction signs of the lanes, such as straight lanes, left turn lanes and / or right turn lanes, but in fact there are more types of traffic signs and their corresponding driving rules in vehicle driving scenarios, such as bus lanes and / or speed limit areas; second, even if some related technologies can recognize traffic signs other than direction signs, they can only recognize the type of traffic signs, but cannot form the detailed rules required for autonomous driving based on them, and match them to the corresponding lanes as a basis for autonomous driving decisions.

[0004] In view of this, there is an urgent need to provide a lane attribute production solution to meet the requirements of autonomous driving scenarios for map data. Summary of the invention

[0005] In order to at least solve one or more of the technical problems mentioned above, the present disclosure proposes a lane attribute production solution in multiple aspects.

[0006] In a first aspect, the present disclosure provides a lane attribute production system including: a map element encoder and a large language model; the output layer of the map element encoder is connected to the input layer of the large language model, the map element encoder is used to process a vectorized map to output a vector encoding result, the vectorized map uses vector features to represent map elements, and the map elements include lanes; the input layer of the large language model also receives at least one of an image encoding result and a text encoding result, and the large language model is configured to: generate lane attribute data based on at least one of the image encoding result and the text encoding result and the vector encoding result, the lane attribute data including: the vehicle's driving rules and the correspondence between the driving rules and the lanes.

[0007] In a second aspect, the present disclosure provides a lane attribute production method, which is applied to a lane attribute production system, the lane attribute production system includes: a map element encoder and a large language model, the output layer of the map element encoder is connected to the input layer of the large language model; the lane attribute production method includes: receiving multimodal data, the multimodal data includes: at least one of an image encoding result and a text encoding result and a vectorized map; using a map element encoder to process the vectorized map to obtain a vector encoding result, wherein the vectorized map uses vector features to represent map elements, and the map elements include lanes; and using a large language model to process at least one of the image encoding result and the text encoding result and the vector encoding result to generate lane attribute data, wherein the lane attribute data includes vehicle driving rules and the correspondence between driving rules and lanes.

[0008] In a third aspect, the present disclosure provides a computer program product, comprising a computer program, which implements the method steps of the second aspect when executed by a processor.

[0009] Through the lane attribute production system provided as above, the disclosed embodiment encodes the map data of the vector modality provided by the vectorized map through a map element encoder, for example, encodes the vector features of the lane, thereby obtaining a vector encoding result, and processes the vector encoding result and the encoding results of other modal data through a large language model with excellent performance in acquiring implicit knowledge. The large language model can process at least one of the image modality and text modality data and the vector modality data, and combine multiple modality data to complete the production of lane attribute data. The lane attribute data not only provides the vehicle's driving rules, but also provides the correspondence between the driving rules and the specific lanes, which can be used to construct an accurate and reliable traffic rule layer in the HD map. The high-definition map with an accurate traffic rule layer can be provided to the autonomous driving and intelligent transportation systems for use, so that the autonomous driving and intelligent transportation systems can make driving decisions that comply with the current lane driving rules on the current lane based on the lane attribute data. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] By reading the detailed description below with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:

[0011] Figure 1 A schematic diagram of the composition structure of an existing high-definition map is shown;

[0012] Figure 2 An exemplary structural diagram of a lane attribute making system according to some embodiments of the present disclosure is shown;

[0013] Figure 3 An exemplary structural diagram showing a lane attribute making system according to some other embodiments of the present disclosure;

[0014] Figure 4 An exemplary workflow diagram of a lane attribute making system according to some other embodiments of the present disclosure is shown;

[0015] Figure 5 An exemplary structural diagram showing a lane attribute making system according to some other embodiments of the present disclosure;

[0016] Figure 6 An exemplary structural diagram of a map element encoder of some embodiments of the present disclosure is shown;

[0017] Figure 7 An exemplary workflow diagram of a map element encoder according to some other embodiments of the present disclosure is shown;

[0018] Figure 8 An exemplary flow chart showing a lane attribute making method according to some embodiments of the present disclosure is shown;

[0019] Fig. 9 An exemplary flow chart of a method for obtaining vector encoding results according to some embodiments of the present disclosure is shown;

[0020] Fig.10 An exemplary flow chart showing a method for producing lane attribute data according to some embodiments of the present disclosure is shown;

[0021] Fig.11 An exemplary flow chart showing a method for training RuleVLM according to some embodiments of the present disclosure is shown;

[0022] Fig.12 An exemplary structural block diagram of an electronic device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.

[0024] It should be understood that the terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0025] It should also be understood that the terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the disclosure. As used in this disclosure and claims, the singular forms of "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" used in this disclosure and claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations.

[0026] As used in this specification and claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.

[0027] The specific implementation of the present disclosure is described in detail below with reference to the accompanying drawings.

[0028] Exemplary application scenarios

[0029] With the rapid development of autonomous vehicles and intelligent transportation systems, the demand for accurate and reliable navigation data is becoming increasingly urgent. High-definition (HD) maps have become an indispensable supporting component for these systems due to their detailed representation of road elements. HD maps can be broken down into the following three core layers: geometry layer, connectivity layer, and traffic rules layer. Figure 1 The schematic diagram of the composition structure of the existing high-definition map is shown in FIG. Figure 1As shown in the figure, the geometry layer provides precise vector data, such as lane dividers and lane center lines, while the connection layer clarifies the relationship between lanes to assist in path planning. The traffic rule layer contains lane-related rule information, such as high-occupancy vehicle lanes, bus lanes, speed limit areas, etc., which can provide information support for the compliance of driving behavior.

[0030] The HD maps currently used in autonomous driving systems have the following defects in the traffic rules layer: First, the traffic rules layer they construct only describes the direction signs of the lanes, such as straight lanes, left turn lanes and / or right turn lanes, but in fact there are more types of traffic signs and their corresponding driving rules in vehicle driving scenarios, such as high-occupancy vehicle lanes, bus lanes and / or speed limit areas; Second, even if some related technologies can identify traffic signs other than direction signs, they can only identify the type of traffic signs, but cannot form detailed rules required for autonomous driving based on them, and match them to the corresponding lanes as the basis for autonomous driving decisions. This means that although the map can provide static structural information of the road, it lacks the ability to dynamically update and interpret traffic rules, cannot provide a structured description consistent with the HD map standard, and cannot support comprehensive autonomous driving applications, which restricts the driving safety of autonomous driving vehicles in complex and changing traffic environments.

[0031] Exemplary Application Scenarios

[0032] In view of this, the disclosed embodiment provides a lane attribute production solution, which processes the vector modality data in the vectorized map through a map element encoder, and provides the processed vector encoding result to a large language model, so that the large language model can generate lane attribute data based on multiple modality data including the vector modality to complete the establishment of the traffic rules layer.

[0033] Figure 2 An exemplary structural diagram of a lane attribute making system according to some embodiments of the present disclosure is shown. Figure 2 As shown, the lane attribute production system includes: a map element encoder and a large language model, wherein the output layer of the map element encoder is connected to the input layer of the large language model, and the map element encoder is used to process the vectorized map to output a vector encoding result. In addition to the vector encoding result, the input layer of the large language model also receives at least one of an image encoding result and a text encoding result, and the large language model is configured to: generate lane attribute data based on at least one of the image encoding result and the text encoding result and the vector encoding result. In the disclosed embodiment, the lane attribute data includes: the driving rules of the vehicle and the corresponding relationship between the driving rules and the lanes.

[0034] According to the above description of the input of the large language model, it can be understood that the large language model shown in the embodiment of the present disclosure can process vector modal data and image modal data, can also process vector modal data and text modal data, and can also process three modal data of vector modal data, image modal data and text modal data, so as to output the detailed rules required for autonomous driving, that is, the lane attribute data in the above embodiment. The lane attribute data not only provides driving rules, but also provides the lane corresponding to a specific driving rule. For example, the lane attribute data describes a driving rule as "the bus lane is not allowed to use the lane between 7:30 and 8:30 in the morning", and the lane attribute data also describes that the lane corresponding to this driving rule is the rightmost lane, then the autonomous driving system can obtain the following information from the lane attribute data: between 7:30 and 8:30 in the morning, the rightmost lane on the map is not allowed to use the lane.

[0035] Large language models (LLMs) are complex machine learning models that are trained to process and generate natural language, which are usually based on deep learning technology and can understand and generate text data. Due to its powerful language understanding ability, large language models can understand the context of input data, not only can they capture the relationship between input data, but also can understand their meaning in a specific context for input data. Further, the disclosed embodiment can use multimodal large language models (MLLMs) as large language models. MLLMs have powerful performance in processing multimodal data and obtaining implicit knowledge of modal alignment. It combines the natural language processing capabilities of large language models with the ability to understand and generate multimodal data, and can establish connections between different modalities, thereby performing tasks that require understanding and generating content across multiple data types, such as: analyzing images based on text descriptions and / or traffic signs to identify corresponding driving rules. As an example, the large language model in the embodiments of the present disclosure can be Qwen-VL, which is an open source large-scale visual language model that can take images and text as input and text as output. It supports analysis tasks in various scenarios such as knowledge question and answer, image caption generation, image question and answer, document question and answer, and fine-grained visual positioning.

[0036] Vectorized map is a kind of map data that uses vector features for map drawing and map element display. It constructs a map by representing multiple map elements in vector form. Map elements may include but are not limited to elements such as lanes and sidewalks. For the map element of lane, it can also be subdivided into sub-elements such as lane centerline and lane separator. Further, a wide range of map elements can be abstracted as a unified point sequence representation. According to geometric features, map elements can be divided into three categories: linear elements, discrete elements and regional elements. Among them, linear elements mainly include lane separators and lane centerlines. By setting sampling points at fixed intervals on these linear elements, the corresponding point sequence representation can be obtained. Taking discrete elements as an example, it can include all elements of regular shapes, such as speed bumps and arrows indicating lane directions. These discrete elements can be represented by the four corner points of their bounding boxes. The order of the four corner points or the order of the corner point sequence can reflect the direction of the element. Taking regional elements as an example, it includes closed-shaped areas, such as sidewalks and detour areas. By setting sampling points at fixed intervals on the boundary of the area, the points on the boundary can be converted into an ordered point sequence. Through the unified point sequence representation, accurate geometric representation can be provided for various map elements on the vectorized map.

[0037] In some embodiments of the present disclosure, the vectorized map provides vector modal data for the model. The map element encoder can use the vectorized map as its input, extract lane-related information from it, form a vector encoding result containing lane information and provide it to the large language model. The large language model can extract lane information features based on the vector encoding result. These lane information features can provide the large language model with information such as lane position, number of lanes and lane distribution. Without obtaining lane information features and only obtaining driving rule features, the large language model can only know that there is a driving rule that needs to be followed by the vehicle within the area represented by the map, but it is impossible to clearly know which lane the vehicle needs to follow when driving on this driving rule. In this case, the traffic rule layer constructed will cause anomalies. For ease of understanding, the following still takes the driving rule of "bus-only lanes are not allowed to use other motor vehicles from 7:30 to 8:30 in the morning" as an example: if this driving rule is matched to the area represented by the entire map, then the automatic driving system will think that all vehicles are not passable from 7:30 to 8:30 in the morning. At this time, the automatic driving system cannot generate a feasible driving route. If the driving rule is randomly matched to a lane, it may result in the driving rule not being matched to the correct lane, and the autonomous driving system may generate a driving route that passes through the bus lane between 7:30 and 8:30 in the morning, resulting in a violation. In other words, the map element encoder is a network structure that processes vector modal data in the lane attribute production system to complete the extraction and integration of lane-related information in the vectorized map.

[0038] It can be seen from the above that the map element encoder provides lane information features for the large language model to form lane attribute data. In order to form lane attribute data, the large language model also needs to obtain driving rule features. In the disclosed embodiment, the driving rule features can be obtained based on image modal data and / or text modal data. In some embodiments, image modal data can be encoded into image encoding results by a visual encoder, and text modal data can be encoded into text encoding results by a text encoder. The large language model can extract driving rule features based on the image encoding results and / or text encoding results. It should be noted that there are many schemes in the prior art for extracting driving rule features from image encoding results and / or text encoding results, such as the visual language model Qwen-VL, which will not be repeated here.

[0039] It can be further understood from the above content that when the large language model performs the action of generating lane attribute data according to at least one of the image encoding results and the text encoding results and the vector encoding results, the large language model is also configured to: extract driving rule features according to the image encoding results and / or the text encoding results; extract lane information features according to the vector encoding results; and match the driving rule features with the lane information features to generate lane attribute data.

[0040] Compared with the existing solutions, the lane attribute production system provided by the disclosed embodiment realizes the processing of vectorized maps by introducing a map element encoder in a large language model, so that the large language model can not only generate separate driving rule features based on image modality data and / or text modality data, but also match the driving rule features to lane information features in combination with vector modality data, and then correspond the driving rules to specific lanes, thereby obtaining complete, accurate, and detailed lane attribute data, which is conducive to the construction of the traffic rule layer of the HD map supporting comprehensive autonomous driving applications. As an example, the large language model can match the driving rule features to the lane information features through the association head prediction technology.

[0041] Relevance head prediction is a technology in the field of computer vision and machine learning, which is usually used in target detection or tracking models to predict features related to the target through a network layer of a relevance head. The relevance head is usually composed of a fully connected layer or a convolutional layer, which receives feature inputs from other parts of the model and outputs features related to the target. In other words, the large language model can use the relevance head to find vector features related to driving rule features for representing lane information, or the large language model can use the relevance head to find driving rule features related to vector features, thereby associating the driving rule features with the vector features representing lane information, thereby outputting lane attribute data including driving rules and their correspondence with lanes on the vectorized map.

[0042] For the convenience of description, we refer to the model combining the map element encoder and the large language model provided in the previous embodiment as RuleVLM, which is designed to process input data of at least two modes including vector mode and generate driving rules corresponding to specific lanes. The lane attribute production system can be understood as a system that uses RuleVLM to complete the construction of the traffic rule layer in the HD map.

[0043] RuleVLM can process vector modal data and image modal data, vector modal data and text modal data, or three modal data, namely, vector modal data, image modal data and text modal data. For the convenience of description, RuleVLM that processes three modalities (image, text and vector) is taken as an example to further explain the structure of the lane attribute production system.

[0044] Figure 3 An exemplary structural diagram of a lane attribute making system according to some other embodiments of the present disclosure is shown. Figure 4 FIG. 2 shows an exemplary workflow diagram of a lane attribute making system according to some other embodiments of the present disclosure. Figure 3 As shown, the lane attribute making system further includes: a visual encoder and a text encoder, wherein the output layer of the visual encoder is connected to the input layer of the large language model, and the output layer of the text encoder is connected to the input layer of the large language model. The visual encoder is used to process the image modality data to output an image encoding result, and the text encoder is used to process the text modality data to output a text encoding result.

[0045] like Figure 4 As shown, the map element encoder, the visual encoder, and the text encoder process the vector modal data, the image modal data, and the text modal data understood by the large language model in turn, so that the three modal data are encoded into a format that can be understood by the large language model, so that the large language model can extract corresponding features from the encoding results, such as driving rule features and lane information features. Specifically, the large language model extracts driving rule features from the encoding results output by the visual encoder and the text encoder, and extracts lane information features from the encoding results output by the map element encoder.

[0046] In order to enable large language models to understand a variety of different modal data, such as Figure 3 and Figure 4As shown, one or more of the map element encoder, the visual encoder, and the text encoder can be connected to the large language model through an adapter to align one or more of the vector encoding results, the image encoding results, and the text encoding results with the embedding space of the large language model. Further, in some embodiments, the adapter can be a linear layer, which is also called a fully connected layer (Fully Connected Layer), which is a basic layer in a neural network. The main function is to perform a linear transformation on the input data and convert the output of the encoder into a form suitable for processing by the large language model. Further, at the beginning of training, the weights of the adapter are usually initialized with identity, so that the initial performance of the model can be guaranteed to be close to the original model. When fine-tuning the target task, the parameters of the map element encoder and the adapter on its output side remain trainable, so that it can be ensured that the weights of the adapter and the map element encoder are adjusted, while the parameters of the visual encoder, the text encoder, and the adapter on its output side remain unchanged, which is conducive to achieving efficient model training.

[0047] In large language models, adapters can achieve efficient parameter fine-tuning. By adding adapters at specific locations in large language models, large language models can quickly adapt to new tasks. Compared with traditional full parameter fine-tuning, adapters only need to train a small number of new parameters, and most parameters of the pre-trained model can be kept unchanged. Because there is no need to adjust the parameters of the entire model, large language models can quickly adapt to new tasks without having to train from scratch, greatly improving analysis efficiency.

[0048] As the model size increases, training can help deal with the complexity of the model and ensure that the model can learn and generalize effectively, thereby improving the performance of the model and making it more accurate and effective on specific tasks. In some embodiments, a large language model can be obtained through LoRA-based training. The full name of LoRA is Low-Rank Adaptation of LLMs, which fine-tunes by optimizing the low-rank decomposition matrix of the changes in the pre-trained model weight matrix while keeping the pre-trained weights unchanged, which greatly reduces the number of trainable parameters in downstream tasks. The core idea of ​​LoRA-based training is to inject a trainable low-rank decomposition matrix into each layer of the Transformer architecture while freezing the pre-trained model weights, thereby reducing the number of training parameters and memory requirements.

[0049] It should be noted that Figure 3 and Figure 4The lane attribute production system shown is equipped with a RuleVLM that processes three modalities (image, text, and vector), and its structure is only an example. In actual application, the RuleVLM carried by the lane attribute production system can simplify one of the visual encoder and the text encoder to form a RuleVLM that processes bimodal data including vector modal data. In other words, in the lane attribute production system shown in the present disclosure, the visual encoder and the text encoder can be set in the RuleVLM or they can be used in combination.

[0050] The large language model in RuleVLM outputs driving rules in text format. In other words, the lane attribute data output by the large language model in RuleVLM is in text format. In order to support comprehensive autonomous driving applications, lane attribute data needs to be converted into a more structured description consistent with the HD map standard. Based on this, some embodiments of the present disclosure also provide a method such as Figure 5 Lane attribute production system, Figure 5 An exemplary structural diagram of a lane attribute making system according to some other embodiments of the present disclosure is shown. Figure 4 and Figure 5 As shown, the RuleVLM it carries also includes: a JSON decoder, the input layer of the JSON decoder is connected to the output layer of the large language model to restore the lane attribute data from text format to JSON format. Data in JSON format is also called formatted data. JSON is a lightweight data exchange format that is easy for users to read and write, and is also easy for machines to parse and generate. Data in JSON format consists of a series of key-value pairs {key: value}, where the key is a string and the value value can be a string, number, Boolean value, array, object, etc. Data in JSON format can be integrated into HD maps for further application.

[0051] The above embodiments introduce the overall network structure of the lane attribute production system. According to the contents described in the above embodiments, the core of the lane attribute production system to realize vector modal data processing lies in the map element encoder. The network structure and working principle of the map element encoder in the lane attribute production system are further described below.

[0052] Figure 6An exemplary structural diagram of a map element encoder of some embodiments of the present disclosure is shown. A map element encoder (MEE) is a functional module in a lane attribute production system that processes data in vector mode provided by a vectorized map, and its structure is similar to a language model. In this embodiment, the vectorized map includes at least vector features for describing lanes, and further, may include vector features for describing lane centerlines and vector features for describing lane separators. In practical applications, as described in the foregoing embodiments, map elements may also include map elements such as speed bumps, arrows indicating lane directions, sidewalks and / or detour areas, and these map elements may also be represented in vector form in the vectorized map.

[0053] like Figure 6 As shown, the map element encoder for receiving and processing a vectorized map includes: an embedding layer, a first vector encoding block and a second vector encoding block, and the embedding layer, the first vector encoding block and the second vector encoding block are connected in sequence.

[0054] The embedding layer is used to receive the vectorized map and extract and convert the vector features therein. The vector features here include at least the vector features of the lanes and may also include the vector features of other map elements. To simplify the description, the vector features mentioned below all represent vector features used to describe map elements, including the vector features of the lanes.

[0055] In this embodiment, the vectorized map is the input of the embedding layer and the input of the map element encoding block. According to the content described in the previous embodiment, in the vectorized map, linear elements such as lane dividers and lane center lines can be represented as point sequences by setting sampling points at fixed intervals, and regional elements such as sidewalks can also be converted into an ordered point sequence by setting sampling points at fixed intervals on the boundaries of the region. It can be seen that the vector features in the vectorized map can be represented as a sequence of points.

[0056] On this basis, the embedding layer can embed the points on each vector feature in the vectorized map to form point embedding. Since each point on the vector feature contains a variety of information, such as the coordinate information of the point, the vector feature to which the point belongs, the relative position between points, etc., the process of embedding the points on the vector feature to obtain point embedding can be divided into multiple processes of extracting different information and forming embedded representations.

[0057] Furthermore, the embedding layer performs embedding conversion on the points on each vector feature in the vectorized map, and can generate vector embedding, type embedding, instance embedding and position embedding that describe each vector feature. The above-mentioned vector embedding, type embedding, instance embedding and position embedding can be understood as subsets of point embedding. Different embeddings are used to represent different types of information. Specifically, vector embedding is used to represent the position of the point on the vector feature, type embedding is used to represent the map element represented by the vector feature where the point is located, instance embedding is used to indicate the vector feature where the point is located, and position embedding is used to represent the relative position between points. In order to facilitate understanding of the difference between type embedding and instance embedding, the two are further explained below: In this embodiment, type embedding is used to distinguish different vector types. For example, the type embedding of the point on the vector feature representing the lane centerline and the point on the vector feature representing the lane separator are different. Instance embedding is used to distinguish different vector instances. In other words, instance embedding is used to distinguish whether different points belong to the same vector feature. The instance embedding of points on the same vector feature is the same, and the instance embedding of points on different vector features is different.

[0058] Since the lane attribute production system provided by the disclosed embodiment is intended to match driving rules to corresponding lanes, the lane attribute production system needs to have the ability to distinguish different lanes. In order to facilitate the large language model in the lane attribute production system to identify different map elements such as lane dividers and lane center lines to complete the distinction between different lanes, this embodiment introduces marker embedding to serve as an index of vector features in the vectorized map, assisting the large language model to understand the map elements and information represented by each vector feature.

[0059] Before completing the assignment of the index of the vector feature, the embedding layer needs to set a blank marker embedding at the beginning of each vector feature in the vectorized map. The marker embedding, vector embedding, type embedding, instance embedding and position embedding are aggregated to form the embedded representation of the vector feature and serve as the input of the first vector coding block. At this time, the blank marker embedding can be understood as a placeholder. First, it reserves space for the subsequent assignment of the first vector coding block and the second vector coding block. Second, for vector features with different lengths of point sequence representation, the input length of the first vector coding block can be unified through the placeholder.

[0060] The first vector encoding block and the second vector encoding block are functional blocks in MEE for encoding vector features. Further, the first vector encoding block includes: M first transformer layers connected in series, where M is a positive integer. Figure 7 An exemplary workflow diagram of a map element encoder of some other embodiments of the present disclosure is shown. Figure 7As shown, the first transformer layer includes an intra-instance attention network (Intra-Instance Attention) and a feed-forward network (FFN) connected in sequence. The first transformer layer, like the traditional transformer architecture, includes multiple stacked encoder layers, each encoder layer includes the following two main parts: a self-attention mechanism and a feed-forward network, and the output of each encoder layer is subjected to residual connection and layer normalization. Different from the traditional transformer architecture, in this embodiment, the self-attention mechanism of the first transformer layer is replaced by an intra-instance attention mechanism to capture the interaction between points within the vector feature. Since the transformer architecture has a multi-layer stacked structure, in order to simplify the expression, Figure 7 Only the simplified structure of the first transformer layer is shown as an example.

[0061] In this embodiment, the first transformer layer can learn and extract information from the embedded representation of the input vector features, including vector embedding, type embedding, instance embedding and position embedding, capture the interaction between the points inside the vector features through the intra-instance attention mechanism, understand the semantic information of each point on the vector features, such as the map element represented by the vector feature where the point is located and the coordinate information of the map element, so as to assign a value to the blank tag embedding.

[0062] Similar to the first vector encoding block, the second vector encoding block also adopts the transformer architecture. The difference is that the second vector encoding block adopts the inter-instance attention mechanism instead of the intra-instance attention mechanism. Specifically, the second vector encoding block includes: N second transformer layers connected in series, N is a positive integer. The second transformer layer adopts the inter-instance attention mechanism to capture the interaction between different vector features. The second transformer layer can also learn and extract information from the embedded representation of the input vector features including vector embedding, type embedding, instance embedding and position embedding, and understand the semantic information different from the first transformer layer, such as the vector features representing different map elements and their relative positions, so as to perform secondary assignment on the marker embedding, form a marker embedding corresponding to the vector feature one by one and use it as the index of each vector feature to obtain a vector encoding result containing lane information. Similar to the first transformer layer, due to the multi-layer stacking structure of the transformer architecture, in order to simplify the expression, Figure 7Only a simplified structure of the second transformer layer is shown as an example, which includes an inter-instance attention mechanism network (Inter-Instance Attention) and a feed-forward network (FFN) connected in sequence.

[0063] According to the above description, through the processing of the first vector encoding block and the second vector encoding block, a marker embedding corresponding to each lane in the vectorized map can be generated. In other words, the marker embedding can be used as an index of the vector feature of the lane. The process of assigning a value to the marker embedding or the process of generating a vector feature index is also the process of mapping the acquired lane information feature to the marker embedding. The lane information obtained here is obtained based on vector embedding, type embedding, instance embedding and position embedding. Vector embedding, type embedding, instance embedding and position embedding respectively provide different lane-related information. For example, the vector embedding of each point on the vector feature of the lane centerline can provide the coordinate information of the lane centerline. Therefore, MEE can output a vector encoding result containing lane information. After obtaining the vector encoding result containing lane information, the large language model can identify the unique vector feature by identifying the index therein. Moreover, since the index is assigned by the first vector encoding block and the second vector encoding block according to vector embedding, type embedding, instance embedding and position embedding, it can pass rich information of the vector feature matching the index to the large language model, thereby facilitating the large language model to extract lane information features and distinguish different lanes, as well as perform feature association analysis through feature association technologies such as association head prediction technology to match driving rules to corresponding lanes.

[0064] In this embodiment, the lane attribute production system introduces MEE to encode vector features in the vectorized map, and uses the marker embedding formed by the encoding as the index of the vector feature to provide the large language model with reference information for distinguishing different lanes and identifying map elements such as lane centerlines. It helps the model to combine data in multiple modes such as images, texts, and vectors to accurately match driving rules and lanes. The lane attribute data formed is conducive to the construction of the traffic rules layer in the HD map.

[0065] Based on the lane attribute making system provided in any of the above embodiments, some embodiments of the present disclosure further implement a lane attribute making method. Figure 8 FIG. 8 is an exemplary flow chart of a lane attribute preparation method 800 according to some embodiments of the present disclosure. Figure 8 As shown, the lane attribute making method includes:

[0066] In step S801, multimodal data is received;

[0067] In step S802, the vectorized map is processed using a map element encoder to obtain a vector encoding result;

[0068] In step S803, at least one of the image encoding result and the text encoding result and the vector encoding result are processed using a large language model to generate lane attribute data.

[0069] In step S801 of the present embodiment, the multimodal data includes a vectorized map, and the vectorized map provides vector modal data. Specifically, the vectorized map provides vector features for describing lanes, for example, vector features for describing lane center lines and lane separators. In addition, the vectorized map may also provide vector features of other map elements. Furthermore, the multimodal data also includes at least one of image encoding results and image encoding results. In other words, the multimodality in the multimodal data refers to vector modality and image modality, or may refer to vector modality and text modality, or may refer to vector modality, image modality, and text modality.

[0070] In step S802 of this embodiment, the internal structure and working principle of the map element encoder can refer to the description of any of the above embodiments, and will not be elaborated here.

[0071] In step S803 of this embodiment, the large language model used is a pre-trained model that can understand a variety of different modal data and analyze them in combination with the multiple modal data to output lane attribute data, which includes the driving rules of the vehicle and the corresponding relationship between the driving rules and the lanes. As an example, the trained large language model can be a trained multimodal large language model. Taking a multimodal large language model that processes three modalities of image, text and vector as an example, after training on a data set containing text, image and vector, the model can extract the required information from different modal data and integrate the extracted information for output.

[0072] In some embodiments, step S803 can be performed as follows: using a large language model to process image encoding results and / or text encoding results to extract driving rule features; using a large language model to process vector encoding results to extract lane information features; and using a large language model to match driving rule features with lane information features to generate lane attribute data.

[0073] In the above process, the trained large language model extracts driving rule features from the image encoding result and / or the text encoding result, and extracts lane information features from the vector encoding result, and matches the two extracted features. The trained large language model can also establish a connection between the extracted different features, that is, establish a connection between the driving rule features and the lane information features in this embodiment, so as to generate lane attribute data.

[0074] As an example, the large language model can establish connections between features through the association head prediction technology, that is, the driving rule features can be associated with the lane information features through the association head prediction to form lane attribute data.

[0075] Relevance head prediction is a technology in the field of computer vision and machine learning, which is usually used in target detection or tracking models to predict features related to the target through a network layer of a relevance head. The relevance head is usually composed of a fully connected layer or a convolutional layer, which receives feature inputs from other parts of the model and outputs features related to the target. In other words, the large language model can use the relevance head to find vector features related to driving rule features, or in other words, the large language model can use the relevance head to find driving rule features related to vector features, thereby associating the driving rule features with the vector features representing lane information, thereby outputting lane attribute data including driving rules and their correspondence with lanes on the vectorized map.

[0076] Based on the lane attribute making method shown in the above embodiment, Fig. 9 FIG. 9 is an exemplary flow chart of a method 900 for obtaining a vector encoding result according to some embodiments of the present disclosure. It can be understood that the lane attribute data obtaining method is a specific implementation of the aforementioned step S802. Figure 8 The features described can be similarly applied here. In this embodiment, the structure of the map element encoder can be seen in Figure 6 , including: an embedding layer, a first vector encoding block and a second vector encoding block connected in sequence. Fig. 9 As shown, the method includes:

[0077] In step S901, the vector features of the lanes in the vectorized map are converted into embedded representations using an embedding layer;

[0078] In step S902, the embedded representation of the vector feature is encoded using the first vector encoding block and the second vector encoding block to form a vector encoding result containing lane information.

[0079] In this embodiment, after receiving the vectorized map, the embedding layer extracts vector features from the vectorized map, and then performs embedding conversion on the points on the vector features of each lane in the vectorized map to obtain vector embedding, type embedding, instance embedding and position embedding. Then, a blank marker embedding is added to the beginning of the vector features of each lane. After that, the marker embedding, vector embedding, type embedding, instance embedding and position embedding are aggregated to form an embedded representation of the vector features of the lane and output it. The embedded representation of the vector features of the lane output by the embedding layer will enter the first vector encoding block and the second vector encoding block. The first vector encoding block and the second vector encoding block encode the embedding of the vector features of the lane to form a vector encoding result containing lane information.

[0080] In this embodiment, the vectorized map is a type of map data that uses vector features for map drawing and map element display. It constructs a map by representing multiple map elements including lane center lines and lane dividers in vector form. According to the previous description of the lane attribute production system, map elements can be abstracted as a unified point sequence representation. According to geometric features, map elements can be divided into three categories: linear elements, discrete elements, and area elements. Taking linear elements as an example, linear elements mainly include lane dividers and lane center lines. By setting sampling points at fixed intervals on these linear elements, the corresponding point sequence representation can be obtained. Through a unified point sequence representation, accurate geometric representation can be provided for various map elements on the vectorized map. Since each point on the vector feature contains a variety of information, such as the coordinate information of the point, the vector feature to which the point belongs, the relative position between points, etc., the process of embedding and transforming the points on the vector feature can be divided into multiple processes of forming different embedded representations based on different information. Therefore, by using the embedding layer in the map element encoder to embed and transform the points on each vector feature in the vectorized map, vector embedding, type embedding, instance embedding and position embedding can be obtained. The specific meaning and difference of each type of embedded representation have been described in detail in the previous embodiments and will not be repeated here.

[0081] The introduction of tag embedding is to facilitate the large language model in the lane attribute production system to distinguish different lanes. In the embedding layer, a blank tag embedding is added as a placeholder at the beginning of each vector feature. First, it reserves space for the subsequent assignment of the first vector encoding block and the second vector encoding block. Second, for vector features with different lengths of point sequence representation, the input length of the first vector encoding block can be unified through the placeholder. Tag embedding, vector embedding, type embedding, instance embedding, and position embedding are aggregated and used as the input of the first vector encoding block in the map element encoder.

[0082] In this embodiment, the first vector encoding block and the second vector encoding block adopt a transformer architecture. The first vector encoding block includes: M first transformer layers connected in series, M is a positive integer, and the first vector encoding block adopts an intra-instance attention mechanism. The second vector encoding block includes: N second transformer layers connected in series, N is a positive integer, and the second vector encoding block adopts an inter-instance attention mechanism. When executing step S902, M first transformer layers connected in series and N second transformer layers connected in series are used to obtain lane information features based on vector embedding, type embedding, instance embedding, and position embedding according to the intra-instance attention mechanism and the inter-instance attention mechanism, and the obtained lane information features are mapped to label embedding to form a vector encoding result containing lane information.

[0083] In the above process, the first vector encoding block and the second vector encoding block map the acquired lane information features to the tag embedding by learning and extracting information from vector embedding, type embedding, instance embedding and position embedding, and assigning values ​​to the blank tag embedding, thereby forming an index corresponding to the vector feature one by one, so as to form a vector encoding result containing the lane information. Specifically, the first vector encoding block uses M first transformer layers connected in series, and assigns values ​​to the blank tag embedding based on vector embedding, type embedding, instance embedding and position embedding by adopting the intra-instance attention mechanism. Then, the second vector encoding block uses N second transformer layers connected in series, and assigns values ​​to the tag embedding based on vector embedding, type embedding, instance embedding and position embedding by adopting the inter-instance attention mechanism. The tag embedding after the secondary assignment is used as the index of each vector feature to obtain the vector encoding result.

[0084] After receiving the vector encoding result output by MEE, the large language model can identify the unique matching vector feature based on the index therein, that is, the assigned tag embedding. For example, it can identify that a certain vector feature is the center line of the leftmost lane, or identify that a certain vector feature is the lane dividing line between the leftmost lane and the middle lane.

[0085] In order to complete the production of lane attribute data, the large language model also needs to obtain driving rule features. As an example, the driving rule features can be obtained from outside the lane attribute production system. At this time, RuleVLM can receive the extracted driving rule features or the encoded image encoding results and / or text encoding results. As another example, the driving rule features can also be obtained from other functional modules inside the lane attribute production system, such as a visual encoder and / or a text encoder. In some embodiments of the present disclosure, the lane attribute production system may also include: a visual encoder and / or a text encoder, wherein the visual encoder can process image modal data to output image encoding results, and the text encoder can process text modal data to output text encoding results. A separate driving rule can be generated based on the image encoding results and the text encoding results. It should be noted that the driving rules at this time have not yet corresponded to specific lanes.

[0086] Combine the following Fig.10 The process of making lane attribute data is explained. According to the description in the above embodiment, in addition to vector modality data, the lane attribute making system can also process image modality data and / or text modality data. At this time, the structure of the lane attribute making system can refer to the structure of the lane attribute making system in conjunction with Figure 3 or Figure 4 The described embodiment includes, in addition to a map element encoder, at least one of a visual encoder and a text encoder, and the multimodal data received by the lane attribute production system may also include at least one of image modality data and text modality data, as well as vector modality data.

[0087] Fig.10 An exemplary flow chart of a lane attribute data preparation method 1000 according to some embodiments of the present disclosure is shown. Fig.10 As shown, the method includes:

[0088] In step S1001, the image modality data is encoded using a visual encoder to output the image encoding result to a large language model;

[0089] In step S1002, the text modality data is encoded using a text encoder to output the text encoding result to the large language model;

[0090] In step S1003, the vector modal data is encoded using a map element encoder to output a vector encoding result to a large language model;

[0091] In step S1004, at least one of the image encoding result and the text encoding result and the vector encoding result are processed using a large language model to generate lane attribute data.

[0092] It should be noted that, in this embodiment, step S1001 and step S1002 can be performed at one or the other. In other words, the lane attribute production system can process image modality data and text modality data at the same time, or only process one of the modality data. Correspondingly, the large language model can process image encoding results and text encoding results at the same time, or only process one of them.

[0093] In addition, it should be noted that there is no strict restriction on the execution sequence of step S1001 to step S1003, and the three steps can be executed in any sequence, for example, in the order of step S1003, step S1002, and step S1001. Alternatively, the three steps of step S1001 to step S1003 can be executed in parallel, and no excessive restrictions are imposed here.

[0094] In this embodiment, step S1003 describes the process of processing the vector modal data using the map element encoder. The specific execution steps of this process can be referred to in conjunction with the above. Fig. 9 The described embodiments will not be described in detail here.

[0095] In this embodiment, the execution process of step S1004 can refer to the above description. Figure 8 Step S803 and related descriptions in the described embodiment are not repeated here. Furthermore, after step S1004 is executed, the lane attribute data output by the large language model is a driving rule in text format. In order to integrate the driving rule into the HD map, it is necessary to convert it into a standard consistent structured description. Therefore, in some embodiments, the lane attribute production system also includes: a JSON decoder, whose input layer is connected to the output layer of the large language model. After the vector encoding result is processed by the large language model, the lane attribute data needs to be restored from the text format to the JSON format using the JSON decoder. Then, the traffic rule layer of the HD map is constructed based on the lane attribute data in JSON format.

[0096] The above embodiments describe the process of completing lane attribute production using a lane attribute production system. Prior to this, the lane attribute production system needs to be constructed and trained through architecture so that it has the corresponding lane attribute production function. The process specifically includes: constructing RuleVLM, and then training RuleVLM to obtain a lane attribute production system. Among them, the process of constructing RuleVLM is as follows: connect the embedding layer, the first vector encoding block and the second vector encoding block in sequence to build a map element encoder, and then connect the output layer of the map element encoder to the input layer of the large language model. Among them, the specific structure of the first vector encoding block and the second vector encoding block has been combined in the previous text. Figure 7The embodiments are described in detail in the embodiments and will not be repeated here.

[0097] In this embodiment, the structure of RuleVLM constructed can be referred to in the above text. Figures 2 to 7 The described embodiments will not be described in detail here.

[0098] Combine the following Fig.11 The training process of RuleVLM is explained. Fig.11 An exemplary flow chart of a RuleVLM training method 1100 according to some embodiments of the present disclosure is shown. Fig.11 As shown, the training method includes:

[0099] In step S1101, a sample data set is obtained;

[0100] In step S1102, the lane attribute data in the JSON format in the sample data set is serialized into a text format to form a sample corpus in the QA format;

[0101] In step S1103, the sample corpus in the QA format is split into a training corpus and an evaluation corpus;

[0102] In step S1104, the RuleVLM is trained using the training corpus;

[0103] In step S1105, the performance of the trained RuleVLM is evaluated using the evaluation corpus until its performance meets the requirements, thereby obtaining a lane attribute production system.

[0104] As an example, the sample dataset used in this embodiment is MapDR, which is a dataset specifically used to extract driving rules from traffic signs and associate them with vectorized high-definition maps. MapDR focuses on complex traffic scenes and collects tens of thousands of video clips covering traffic scenes in multiple Chinese cities including Beijing and Shanghai, each of which contains at least one traffic sign. In addition to the vectorized representation of dividers, boundaries, and center lines, MapDR also provides structural annotations of traffic regulations and their association with lanes.

[0105] It should be noted that what is introduced above is an exemplary data set that can be selected in the embodiment of the present disclosure. In actual application, there may be other data sets of the same type that are suitable for the training method shown in the present disclosure, and no excessive restrictions are made here.

[0106] Since a standard and consistent structured description is required when integrating driving rules into HD maps in order to support comprehensive autonomous driving applications, the lane attribute data in the sample dataset is usually represented in JSON format. However, compared with the QA corpus, its training effect on models that train language processing tasks is slightly worse. The reason is that the QA corpus usually contains richer language phenomena and more complex contexts, which is more suitable for models that train language processing tasks. In contrast, although JSON data is structured, it is not as diverse and complex as the QA corpus. In addition, the training of QA corpus is more task-oriented. The model can learn how to provide answers to specific questions through the QA corpus, which can improve the performance of the model on specific tasks. Moreover, through the QA corpus, the model can learn how to extract information from a given text, which has a positive impact on subsequent fine-tuning and the execution of specific tasks.

[0107] Based on the above reasons, after obtaining the sample data set, this embodiment will also serialize the lane attribute data in JSON format into text format to form a sample corpus in QA format. In addition, in addition to being used as training data for the model to learn, the sample data set can also be used as evaluation data to evaluate the performance of the trained model to reflect whether the model has been trained or whether it needs to be further adjusted. As an example, the trained RuleVLM is tested using the evaluation corpus, and the recall rate of the test results is calculated. When the recall rate meets the standard, it is determined that the RuleVLM has been trained. Otherwise, the model is fine-tuned for a second time using the training corpus until its recall rate meets the standard. As an example, the recall rate threshold can be set to 80% or other values, and the specific value can be set according to actual needs. There are no excessive restrictions here.

[0108] It should be noted that, in some embodiments, the lane attribute data in JSON format may be first serialized into a sample corpus in QA format, and then split into a training corpus for training and an evaluation corpus for evaluation, as described in the previous embodiments. In other embodiments, the sample data set may be first split to obtain a training data set and an evaluation data set, and then the training data set and the evaluation data set may be format converted respectively. The splitting methods of the above two data sets are applicable to the present disclosure, and no excessive restrictions are imposed here. In addition, the sample data set may be split according to a certain ratio, for example, the training data and the evaluation data may be split according to a ratio of 9:1, 8:2 or other ratios. There are no strict requirements for the splitting ratio, and it can be adjusted according to actual needs.

[0109] Furthermore, when using the training corpus to train RuleVLM, the LoRA-based training method can be used to train the large language model therein. Furthermore, in RuleVLM, the map element encoder can be connected to the large language model through an adapter, and the adapter can be a linear layer. Then, when executing step S1104, the parameters of the map element encoder and its adapter can be kept in a trainable state, and the parameters of the visual encoder and the text encoder and their adapters can be fixed. Then, the trainable parameters in RuleVLM are trained according to the training corpus, thereby achieving efficient parameter fine-tuning.

[0110] The above embodiments describe how to optimize the training of RuleVLM by format conversion and / or selecting an efficient training method. In other embodiments, in order to optimize the training of RuleVLM to obtain a RuleVLM with enhanced performance, it is also necessary to prevent overfitting from occurring during the training process.

[0111] To address the overfitting problem, some embodiments of the present disclosure process the sample data set as follows, including: obtaining the original sample data set, and then sequentially adjusting the lane centerline data of the vector mode in the original sample data set to form an anti-overfitting sample data set. This method improves the diversity and amount of training data, and can achieve a certain degree of anti-overfitting.

[0112] In summary, the disclosed embodiment provides a lane attribute production system, which can be equipped with a RuleVLM including a map element encoder and a large language model, and defines a task of integrating traffic rules into a vectorized high-definition map for execution by the RuleVLM. The RuleVLM completes the processing of the vector modal data of the vectorized map through the introduced MEE, and analyzes multiple modal data including the vector modality and other modalities through the large language model, and then outputs the lane attribute data that matches the driving rules to the corresponding lane, so as to facilitate the construction of the traffic rule layer in the HD map.

[0113] The disclosed embodiment also provides a lane attribute production method, which performs the task of integrating traffic rules into vectorized high-definition maps through RuleVLM. Specifically, RuleVLM can combine data of multiple modalities to associate driving rule features with vector features in the vectorized map, thereby completing driving rule-lane matching. With RuleVLM, the method can generate detailed rule descriptions required for autonomous driving, and the rule descriptions can also be matched with specific lanes, which is conducive to building an accurate traffic rule layer, thereby forming a reliable HD map to support comprehensive autonomous driving applications.

[0114] In order to implement the method steps described in the foregoing text of this disclosure in conjunction with the accompanying drawings at the software and hardware level, the present disclosure also provides the following Fig.12 The electronic device shown. Specifically, Fig.12 An exemplary structural block diagram of an electronic device 1200 according to an embodiment of the present disclosure is shown.

[0115] like Fig.12 As shown, the electronic device 1200 disclosed in the present invention may include a processor 1210 and a memory 1220. Specifically, the memory 1220 stores executable program instructions. When the program instructions are executed by the processor 1210, the electronic device implements the above-mentioned Figure 8-Figure 11 The method steps are described.

[0116] It is understood that in order to clearly illustrate the solution of the present disclosure and avoid confusion with the prior art, Fig.12 The electronic device 1200 of the present disclosure only shows the components related to the embodiment of the present disclosure, and omits the components that may be necessary for implementing the embodiment of the present disclosure but belong to the scope of the prior art. Therefore, based on the content disclosed in the present disclosure, a person skilled in the art can clearly understand that the electronic device 1200 of the present disclosure can also include components related to the embodiment of the present disclosure. Fig.12 The constituent elements shown in are different from the common constituent elements.

[0117] In an exemplary implementation scenario, the processor 1210 described above can control the overall operation of the electronic device 1200. For example, the processor 1210 can control the operation of the electronic device 1200 by executing the program stored in the memory 1220. In terms of implementation, the processor 1210 disclosed herein can be implemented by a central processing unit (CPU), an application processor (AP), an artificial intelligence processor chip (IPU), etc. provided in the electronic device 1200. Further, the processor 1210 disclosed herein can also be implemented in any appropriate manner. For example, the processor 1210 can take the form of a computer-readable medium, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc., such as a microprocessor or a processor and a computer-readable program code (such as software or firmware) that can be executed by the (micro) processor.

[0118] In terms of storage content, the memory 1220 can be used to store hardware of various data and instructions processed in the electronic device 1200. For example, the memory 1220 can store processed data and data to be processed in the electronic device 1200. The memory 1220 can store data sets that have been processed or to be processed by the processor 1210. In addition, the memory 1220 can store applications, drivers, etc. to be driven by the electronic device 1200. For example: the memory 1220 can store various programs to be executed by the processor 1210. The memory 1220 can be a DRAM, but the present disclosure is not limited thereto. In terms of type, the memory 1220 may include at least one of a volatile memory or a non-volatile memory. The non-volatile memory may include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a flash memory, a phase change RAM (PRAM), a magnetic RAM (MRAM), a resistive RAM (RRAM), a ferroelectric RAM (FRAM), and the like. The volatile memory may include dynamic RAM (DRAM), static RAM (SRAM), synchronous DRAM (SDRAM), PRAM, MRAM, RRAM, ferroelectric RAM (FeRAM), etc. In an embodiment, the memory 1220 may include at least one of a hard disk drive (HDD), a solid state drive (SSD), a high-density flash memory (CF), a secure digital (SD) card, a micro secure digital (Micro-SD) card, a mini secure digital (Min i-SD) card, an extreme digital (xD) card, caches, or a memory stick.

[0119] In summary, the specific functions implemented by the memory 1220 and the processor 1210 of the electronic device 1200 provided in the implementation manner of this specification can be interpreted in comparison with the aforementioned implementation manner in this specification, and can achieve the technical effects of the aforementioned implementation manner, and will not be repeated here.

[0120] Additionally or optionally, the present disclosure may also be implemented as a non-temporary machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which computer program instructions (or computer program, or computer instruction code) are stored. When the computer program instructions (or computer program, or computer instruction code) are executed by a processor of an electronic device (or electronic device, server, etc.), the processor executes part or all of the steps of the above-mentioned method according to the present disclosure.

[0121] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may think of many changes, modifications, and alternatives without departing from the thought and spirit of the present disclosure. It should be understood that in the process of practicing the present disclosure, various alternatives to the embodiments of the present disclosure described herein may be adopted. The attached claims are intended to define the scope of protection of the present disclosure, and therefore cover equivalents or alternatives within the scope of these claims.

[0122] The collection and acquisition of various data in this disclosure complies with relevant laws and regulations and is authorized by the data provider. Any organization or individual that needs to obtain external data must obtain authorization in accordance with the law and ensure data security. It is prohibited to illegally collect, use, process, or transmit unauthorized or unprotected data, or to illegally buy, sell, provide, or disclose unauthorized or unprotected data.

Claims

1. A lane attribute making system, characterized in that: include: Map element encoders and large language models; The output layer of the map element encoder is connected to the input layer of the large language model, the map element encoder is used to process the vectorized map to output a vector encoding result, the vectorized map uses vector features to represent map elements, and the map elements include lanes; The input layer of the large language model further receives at least one of an image encoding result and a text encoding result, and the large language model is configured to: Lane attribute data is generated according to at least one of the image encoding result and the text encoding result and the vector encoding result. The lane attribute data includes: driving rules of the vehicle and the corresponding relationship between the driving rules and the lane.

2. The lane attribute making system according to claim 1, characterized in that: The large language model is also configured to: Extracting driving rule features according to the image encoding result and / or the text encoding result; Extracting lane information features according to the vector encoding result; as well as The driving rule feature is matched with the lane information feature to generate lane attribute data.

3. The lane attribute making system according to claim 1 or 2, characterized in that: The map element encoder includes: an embedding layer, a first vector encoding block and a second vector encoding block connected in sequence, the embedding layer is used to convert the vector features of the lanes in the vectorized map from vector representation to embedded representation, and the first vector encoding block and the second vector encoding block are used to encode the vector features of the embedded representation into a vector encoding result containing lane information.

4. The lane attribute making system according to claim 3, characterized in that: The embedding layer is used to perform embedding transformation on points on the vector features of each lane in the vectorized map to generate vector embedding, type embedding, instance embedding and position embedding describing the vector features of each lane, and marker embedding for setting a blank at the beginning of the vector features of each lane in the vectorized map, wherein the marker embedding, the vector embedding, the type embedding, the instance embedding and the position embedding are aggregated to form an embedded representation of the vector features of the lane; Among them, the vector embedding is used to represent the position of a point on a vector feature, the type embedding is used to represent the map element represented by the vector feature where the point is located, the instance embedding is used to indicate the vector feature where the point is located, the position embedding is used to represent the relative position between points, and the tag embedding is used as an index for each vector feature.

5. The lane attribute making system according to claim 4, characterized in that: The first vector encoding block includes: M first transformer layers connected in series, where M is a positive integer, the first transformer layers adopt an intra-instance attention mechanism, and the M first transformer layers connected in series are used to receive an embedded representation of a vector feature of a lane, and assign a blank mark embedding according to the vector embedding, the type embedding, the instance embedding, and the position embedding; The second vector encoding block includes: N second transformer layers connected in series, N is a positive integer, the second transformer layers adopt an inter-instance attention mechanism, and the N second transformer layers connected in series are used to perform secondary assignment on the label embedding according to the vector embedding, the type embedding, the instance embedding and the position embedding to form a label embedding corresponding to the vector feature one by one and use it as the index of each vector feature to obtain a vector encoding result containing lane information.

6. The lane attribute making system according to claim 1, characterized in that: The lane attribute data output by the large language model is driving rules in text format; The lane attribute making system further includes: a JSON decoder, the input layer of which is connected to the output layer of the large language model to restore the lane attribute data from a text format to a JSON format.

7. A lane attribute making method, characterized in that: It is applied to a lane attribute production system, the lane attribute production system comprises: a map element encoder and a large language model, the output layer of the map element encoder is connected to the input layer of the large language model; the lane attribute production method comprises: Receiving multimodal data, the multimodal data comprising: at least one of an image encoding result and a text encoding result and a vectorized map; Processing the vectorized map using the map element encoder to obtain a vector encoding result, wherein the vectorized map uses vector features to represent map elements, and the map elements include lanes; and The large language model is used to process at least one of the image encoding result and the text encoding result and the vector encoding result to generate lane attribute data, wherein the lane attribute data includes driving rules of the vehicle and the corresponding relationship between the driving rules and the lane.

8. The lane attribute making method according to claim 7, characterized in that: The method further comprises: using the large language model to process at least one of the image encoding result and the text encoding result and the vector encoding result to generate lane attribute data, including: Processing the image encoding result and / or the text encoding result using the large language model to extract driving rule features; Processing the vector encoding result using the large language model to extract lane information features; and The driving rule feature is matched with the lane information feature using the large language model to generate the lane attribute data.

9. The lane attribute making method according to claim 7, characterized in that: The map element encoder comprises: an embedding layer, a first vector encoding block and a second vector encoding block connected in sequence; wherein processing the vectorized map by using the map element encoder to obtain a vector encoding result containing lane information comprises: Using the embedding layer to convert the vector features of the lanes in the vectorized map into an embedded representation; and The embedded representation of the vector feature is encoded using the first vector encoding block and the second vector encoding block to form a vector encoding result containing lane information.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 7 to 9 are implemented.

Citation Information

Patent Citations

  • Top-down scene prediction based on motion data

    CN114245885A

  • Method and device for training model and detecting road surface element change

    CN117555911A

  • Method for identifying rule packet for generating current traffic situation of motor vehicle from traffic rules on basis of traffic signs

    CN117593894A

  • Map element detection method and device, model training method and device and electronic equipment

    CN117611547A

  • Map generation method and device, equipment and storage medium

    CN117928574A

Cited By

  • Lane attribute creation system and creation method, and computer program product

    WO2026137952A1