Method for constructing map based on large model, vehicle control method, apparatus, electronic device, storage medium, and computer program

The map construction method using a large-scale model addresses the challenges of low lane accuracy and inefficient updates by processing detection target images and target presentation information, resulting in improved autonomous vehicle navigation.

JP2025094094AActive Publication Date: 2025-06-24BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025043860
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-16
Filing Date
2025-03-18
Publication Date
2025-06-24
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

Existing map construction methods for autonomous driving face challenges such as low lane accuracy, inefficient map updates, and difficulty in accurately representing lane attributes like boundary lines and directions.

Method used

A map construction method based on a large-scale model that processes detection target images and target presentation information to generate an area road map, using a large-scale model to improve lane detection accuracy and map updating efficiency.

Benefits of technology

The method enhances lane attribute detection accuracy and map updating efficiency, improving the safety and efficiency of autonomous vehicle navigation by providing a more accurate and timely road map.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025094094000001_ABST
    Figure 2025094094000001_ABST
Patent Text Reader

Abstract

To provide a method for constructing a map based on a large model, a vehicle control method, an apparatus, an electronic device, a storage medium, and a computer program which are applicable to scenes of automatic driving, unmanned driving, etc.SOLUTION: A method for constructing a map based on a large model, includes acquiring an associated-area lane attribute and an image to be detected which is collected by a vehicle-side sensor. The image to be detected represents a road area to be detected. The associated-area lane attribute corresponds to an associated road area. The associated road area and the road area to be detected meet a predetermined similarity condition. The method further includes constructing target prompt information based on the associated-area lane attribute, and processing the target prompt information and the image to be detected by using the large model to obtain an area road map for the road area to be detected.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to the fields of computer vision, deep learning, large-scale models, and generative model technologies, and is applicable to scenarios such as autonomous driving and driverless driving. Specifically, it relates to a map construction method, a vehicle control method, an apparatus, an electronic device, a storage medium, and a computer program based on a large-scale model.

Background Art

[0002] With the rapid development of science and technology, the number of vehicles running on the road shows a relatively rapid increasing trend. Vehicles can realize the autonomous driving function through the road map, and drivers can also assist in the driving of the vehicle based on a relatively accurate road map, improving driving safety.

Summary of the Invention

Problems to be Solved by the Invention

[0003] The present disclosure provides a map construction method, a vehicle control method, an apparatus, an electronic device, a storage medium, and a computer program based on a large-scale model.

Means for Solving the Problems

[0004] According to one aspect of the present disclosure, it is to obtain a detection target image collected by a related area lane attribute and a vehicle-side sensor, where the detection target image represents a detection target road area, the related area lane attribute corresponds to a related road area, and a predetermined similarity condition is satisfied between the related road area and the detection target road area; construct target presentation information based on the related area lane attribute; and use a large-scale model to process the target presentation information and the detection target image to obtain an area road map of the detection target road area. A map construction method based on a large-scale model is provided.

[0005] According to another aspect of the present disclosure, there is provided a vehicle control method including controlling the running of a vehicle based on a road map constructed by the large-scale model-based map construction method provided by this embodiment.

[0006] According to another aspect of the present disclosure, there is provided an acquisition module that acquires a detection target image collected by a related area lane attribute and a vehicle-side sensor, the detection target image represents a detection target road area, the related area lane attribute corresponds to a related road area, and a predetermined similarity condition is satisfied between the related road area and the detection target road area; a target presentation information construction module that constructs target presentation information based on the related area lane attribute; and a first map construction module that processes the target presentation information and the detection target image using a large-scale model to obtain an area road map of the detection target road area. There is provided a map construction apparatus based on a large-scale model including these modules.

[0007] According to another aspect of the present disclosure, there is provided a vehicle control apparatus including a vehicle control module that controls the running of a vehicle based on a road map constructed by the large-scale model-based map construction apparatus provided by this embodiment.

[0008] According to an embodiment of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the large-scale model-based map construction method provided by the embodiment of the present disclosure.

[0009] According to an embodiment of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the vehicle control method provided by the embodiment of the present disclosure.

[0010] According to an embodiment of the present disclosure, an autonomous vehicle is provided, including an electronic device for executing a vehicle control method provided by the embodiment of the present disclosure.

[0011] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method provided by the embodiment of the present disclosure, is provided.

[0012] According to an embodiment of the present disclosure, a computer program is provided which, when executed by a processor, implements the method provided by the embodiment of the present disclosure.

[0013] It should be understood that the content described in this part is not intended to indicate the key points or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be easily understood from the following description.

Brief Description of the Drawings

[0014] The drawings are for better understanding of the present invention and do not limit the present disclosure.

[0015]

Figure 1

[0016]

Figure 2

[0017]

Figure 3

[0018]

Figure 4

[0019]

Figure 5

[0020]

Figure 6

[0021]

Figure 7

[0022]

Figure 8

[0023]

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0024] Hereinafter, exemplary embodiments of the present disclosure will be described with reference to the drawings. Here, various details of the embodiments of the present disclosure are included for easier understanding, and they should be considered exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clear and concise description, descriptions of well-known functions and configurations are omitted in the following description.

[0025] In the technical solution of the present disclosure, the acquisition, storage, application, etc. of the relevant user personal information all comply with the provisions of relevant laws and regulations, take necessary security measures, and do not violate public order and good customs.

[0026] A road map with a lane-level navigation function can improve the safety and traffic efficiency of vehicle automatic driving. However, the inventor has discovered that in the construction method of the road map, the lane accuracy may be low, the update efficiency of the road map is low, and it may be difficult to accurately represent lane attributes such as the actual lane boundary line type and the lane boundary line direction.

[0027] The embodiments of the present disclosure provide a map construction method, a vehicle control method, an apparatus, an electronic device, a storage medium, and a computer program based on a large-scale model. The map construction method based on the large-scale model includes: obtaining a detection target image collected by a relevant area lane attribute and a vehicle-side sensor, where the detection target image represents a detection target road area, the relevant area lane attribute corresponds to a relevant road area, and a predetermined similarity condition is satisfied between the relevant road area and the detection target road area; constructing target presentation information based on the relevant area lane attribute; and using the large-scale model to process the target presentation information and the detection target image to obtain an area road map of the detection target road area.

[0028] According to an embodiment of the present disclosure, by acquiring a detection target image collected by a vehicle-side sensor, obtaining a related area lane attribute of a related road area that satisfies a predetermined similarity condition with a detection target road area, and constructing target presentation information based on the related area lane attribute, lane attribute information highly relevant to the lane of the detection target road area can be included in the target presentation information. Furthermore, by processing the target presentation information and the detection target image using a large-scale model, it is possible to realize using the lane attribute information represented by the target presentation information as auxiliary understanding content of the detection target image. Under the condition of controlling the large-scale model based on the target presentation information to relatively well understand the lane attributes highly relevant to the detection target road area, the large-scale model can be controlled to relatively accurately detect the lanes of the detection target road area, and furthermore, the detection accuracy and drawing accuracy of the area road map can be improved. At the same time, a detection target image can also be acquired by a vehicle-side sensor of a vehicle traveling near the detection target road area. Furthermore, a large number of detection target images for constructing an area road map can be acquired relatively conveniently, improving the convenience and timeliness of area road map updating.

[0029] FIG. 1 schematically shows an exemplary system architecture to which a map construction method and apparatus based on a large-scale model according to an embodiment of the present disclosure can be applied.

[0030] Note that what is shown in FIG. 1 is merely an example of a system architecture to which an embodiment of the present disclosure can be applied for those skilled in the art to understand the technical content of the present disclosure, and it does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments or scenes. For example, in another embodiment, an exemplary system architecture to which a map construction method and apparatus based on a large-scale model can be applied may include a terminal device, but the terminal device can implement the map construction method and apparatus based on the large-scale model provided by the embodiments of the present disclosure without interacting with the server.

[0031] As shown in FIG. 1, the system architecture 100 according to the embodiment may include a first terminal device 101, a second terminal device 102, a vehicle 103, a network 104, and a server 105. The network 104 provides a medium for communication links among the first terminal device 101, the second terminal device 102, and the server 105. The network 104 may include various connection types such as wired and / or wireless communication links.

[0032] The user can use the first terminal device 101 and the second terminal device 102 to interact with the server 105 via the network 104 to send and receive messages and the like. Various communication client applications such as a knowledge browsing system application, a web browser application, a search system application, an instant messaging tool, a mailbox client, and / or social platform software may be installed on the first terminal device 101 and the second terminal device 102 (for illustration only).

[0033] The first terminal device 101 and the second terminal device 102 may be various electronic devices having a display and supporting web page browsing, including but not limited to smartphones, tablets, laptop computers, and desktop computers.

[0034] A vehicle-side sensor may be attached to the vehicle 103. The vehicle-side sensor may include, but is not limited to, an image sensor, and may be other types of sensors such as a lidar and a millimeter-wave radar. The vehicle 103 may be any type of vehicle such as a passenger car or a truck. A communication device may be installed in the vehicle 103, and the vehicle 103 may perform information transmission with the server 105 via the network 104 based on the communication device.

[0035] Server 105 may be a server that provides various services. For example, it may be a background management server (for illustration only) that supports the content browsed by the user using the first terminal device 101 and the second terminal device 102. The background management server can perform processing such as analysis on the received data such as user requests, and feedback the processing results (for example, web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0036] The server may also be called a cloud computing server or a cloud host, and may be a cloud server which is one of the host products in a cloud computing service system. It solves the defects of the large management difficulty and weak service scalability existing in the conventional physical host and VPS service (abbreviated as "Virtual Private Server" or "VPS"). The server may be a server of a distributed system, or a server combined with a blockchain.

[0037] It should be noted that the map construction method based on the large-scale model provided by the embodiments of the present disclosure may generally be executed by the server 105. Correspondingly, the map construction device based on the large-scale model provided by the embodiments of the present disclosure may generally be provided in the server 105. Different from the server 105, the map construction method based on the large-scale model provided by the embodiments of the present disclosure may also be executed by a server or a server cluster that can communicate with the first terminal device 101, the second terminal device 102, the vehicle 103, and / or the server 105. Correspondingly, different from the server 105, the map construction device based on the large-scale model provided by the embodiments of the present disclosure may also be provided in a server or a server cluster that can communicate with the first terminal device 101, the second terminal device 102, the vehicle 103, and / or the server 105.

[0038] It should be understood that the numbers of terminal devices, vehicles, networks, and servers in FIG. 1 are merely exemplary. Any number of terminal devices, vehicles, networks, and servers may be provided according to the needs of implementation.

[0039] Figure 2 schematically shows a flowchart of a map construction method based on a large-scale model according to an embodiment of the present disclosure.

[0040] As shown in Figure 2, the map construction method based on the large-scale model includes operations S210 to S240.

[0041] In operation S210, detection target images collected by the relevant area lane attribute and the vehicle-side sensor are acquired.

[0042] In operation S220, target presentation information is constructed based on the relevant area lane attribute.

[0043] In operation S230, the large-scale model is used to process the target presentation information and the detection target images to obtain an area road map of the detection target road area.

[0044] According to an embodiment of the present disclosure, the detection target image can represent the detection target road area, and for example, it may be a single-frame or multi-frame image obtained by collecting an image of the detection target road area. The vehicle-side sensor may include any type of image collection device attached to the vehicle, such as a monocular camera, a surround view camera, a depth camera, etc. The detection target image may include any type of image, such as a color image, a grayscale image, etc.

[0045] According to an embodiment of the present disclosure, the lane attribute of the related area corresponds to the related road area, and the related road area and the detected target road area satisfy a predetermined similarity condition. The related road area may include an area that at least partially overlaps with the detected target road area, or the related road area may further include a road area adjacent to the detected target road area, for example, a road area adjacent to the detected target road area or a road area within a predetermined distance range from the detected target road area. Alternatively, the related road area may further include an area having a predetermined traffic relationship with the detected target road area. For example, the detected target road area and the related road area belong to the same traffic link (for example, the same highway, the same elevated traffic route). The embodiment of the present disclosure does not limit the specific setting method of the predetermined similarity condition, and can be designed according to actual needs, as long as it can meet the actual needs.

[0046] According to an embodiment of the present disclosure, the lane attribute of the related area may include the attributes of the lane boundary lines in the related road area, such as lane shape, lane type, position, lane boundary line topological relationship, etc. Note that the lane type may include the traffic direction indicated by the lane boundary line (for example, left turn, straight ahead, etc.), or the traffic rules indicated by the lane boundary line, such as vehicle cut-in permission, oncoming traffic indication, etc.

[0047] According to an embodiment of the present disclosure, the target presentation information may be a serialized identifier or a tensor that helps the large-scale model understand the detection task. Based on the included lane attributes of the related area, the target presentation information assists the large-scale model to detect the lane attributes in the detected target image relatively accurately, avoiding obvious attribute conflicts (such as changes in the direction indicated by the lane boundary line) between the lane attributes in the area lane detection result and the lane attributes of the related area, and improving the detection accuracy for the lane boundary lines in the detected target road area.

[0048] According to an embodiment of the present disclosure, a large-scale model may be a model having a large number of parameters in its model structure, and the order of the parameters is generally on the order of tens of millions, hundreds of millions or more, and may reach the order of billions or tens of billions. For the network structure of the large-scale model, for example, a network structure such as UFO (Unified Featuer Optimization) can be adopted. By processing the detection target image and the target presentation information with the large-scale model, the strong feature understanding ability and expression ability of the large-scale model can be exerted, the detection target image with a large data scale can be efficiently processed, and the detection accuracy and detection efficiency of the area road map can be improved.

[0049] According to an embodiment of the present disclosure, the related area lane attribute includes the related area lane position and the related area lane type. The related area lane position may include the coordinate position of the lane boundary line. The related area lane type may include any type such as the color of the lane boundary line and the traffic rule indication type in the related area.

[0050] In one example, the related area lane position may be the coordinates of one or more points of the lane boundary line in the related area.

[0051] According to an embodiment of the present disclosure, constructing the target presentation information based on the related area lane attribute may include performing feature fusion on the related area lane position and the related area lane type to obtain the related lane position and type features, and constructing the target presentation information based on the related lane position and type features.

[0052] According to an embodiment of the present disclosure, performing feature fusion on the related area lane position and the related area lane type may include processing the related area lane position and the related area lane type based on a fusion algorithm. The fusion algorithm may include, for example, matrix multiplication, concatenation, an attention algorithm, etc. The embodiments of the present disclosure do not limit the specific algorithm type of the fusion algorithm.

[0053] Note that, for the related area lane position and the related area lane type, processes such as tokenization and embedding can be performed to obtain a feature vector representing the related area lane position and the related area lane type. Thereafter, feature fusion is performed on the feature vector.

[0054] According to an embodiment of the present disclosure, constructing the target presentation information based on the related lane position and type features may further include constructing the target presentation information based on the related lane position and type features and features representing other related area lane attributes.

[0055] In one example, the target presentation information may include the related lane spatial features obtained by performing spatial feature extraction on the related area lane attributes and the related lane position and type features. The related area lane attributes can be represented based on an SD map (Standard Definition Map), and the SD map can be processed based on any type of neural network algorithm such as a convolutional neural network to obtain the related lane spatial features.

[0056] In one example, by processing the related area lane attributes based on FPN (Feature Pyramid Networks), multiple hierarchical feature extractions and fusions are performed on the spatial features of the related area lane attributes, and the obtained related lane spatial features can relatively accurately represent the spatial semantic information of the lane boundary lines in the related road area.

[0057] In one example, constructing the target presentation information based on the related lane position and type features may include using the related lane position and type features as the target presentation information.

[0058] According to an embodiment of the present disclosure, a plurality of related road areas may be included, and the related area lane attributes of the plurality of related road areas correspond to the plurality of related lane positions and type features.

[0059] According to the embodiments of the present disclosure, the plurality of related road areas and the road area to be detected can meet a predetermined similarity condition. For example, they can meet a predetermined distance condition, a predetermined traffic relationship condition, etc., and the embodiments of the present disclosure will omit the description here.

[0060] According to the embodiments of the present disclosure, the plurality of related lane positions and type features may include at least one reference position and type feature, and the reference position and type feature may be the related lane positions and type features obtained by executing the method provided by the embodiments of the present disclosure based on other images to be detected. For example, in the process of executing the method provided by the embodiments of the present disclosure multiple times, in the current i-th execution process, the current i-th related lane positions and type features can be obtained, and the related lane positions and type features obtained from the first time to the i-1-th time can be used as the reference position and type features.

[0061] According to the embodiments of the present disclosure, constructing the target presentation information based on the related lane positions and type features may include specifying the relevance weight of each of the plurality of related lane positions and type features based on at least one related lane position and type feature, and fusing the plurality of related lane positions and type features based on the plurality of relevance weights to obtain the target presentation information.

[0062] According to the embodiments of the present disclosure, the relevance weight can represent the degree of relevance between the related road area and the road area to be detected. For example, the degree of relevance can represent the area overlap between the related road area and the road area to be detected. However, it is not limited thereto, and the degree of relevance can also represent the distance between the separated related road area and the road area to be detected. Also, for example, the degree of relevance can represent other related relationships such as the road traffic relationship between the related road area and the road area to be detected. The road traffic relationship may indicate relationships such as convergence, divergence, and connection between lane boundaries, for example.

[0063] Note that the relevance weight can be represented based on any method such as a matrix, numerical value, etc., and the embodiments of the present disclosure do not limit this.

[0064] According to an embodiment of the present disclosure, fusing a plurality of relevant lane positions and type features based on a plurality of relevance weights may include performing a weighted average calculation based on the respective relevance weights of the plurality of relevant lane positions and type features to obtain a fused relevant area fusion feature. The target presentation information is specified based on the fused relevant area fusion feature.

[0065] According to an embodiment of the present disclosure, fusing a plurality of relevant lane positions and type features based on a plurality of relevance weights may further include multiplying the relevance weights by the relevant lane positions and type features to obtain a plurality of initial relevant area fusion features. The plurality of initial relevant area fusion features are processed based on an attention network algorithm or another type of neural network algorithm to obtain a fused relevant area fusion feature. The target presentation information is specified based on the fused relevant area fusion feature.

[0066] According to an embodiment of the present disclosure, specifying the target presentation information based on the fused relevant area fusion feature may further include constructing the target presentation information based on the relevant area fusion feature and the feature representing other relevant area lane attributes.

[0067] In one example, the target presentation information can be constructed based on the relevant area fusion feature and the relevant lane space feature. The target presentation information may include the relevant area fusion feature and the relevant lane space feature.

[0068] According to an embodiment of the present disclosure, the large-scale model may include an encoder and a decoder. The encoder and the decoder can be constructed based on an attention network algorithm. For example, the encoder and the decoder can be constructed based on a self-attention network algorithm.

[0069] According to an embodiment of the present disclosure, processing the target presentation information and the detection target image using a large-scale model to obtain an area road map of the detection target road area may include fusing the detection target feature and the target presentation information using an encoder to obtain a detection target area feature in a bird's-eye view space, and based on an attention mechanism, processing the target presentation information and the detection target area feature using a decoder to obtain an area road map.

[0070] According to an embodiment of the present disclosure, the detection target feature may be specified based on the detection target image. For example, the detection target image can be processed based on any type of network algorithm such as a convolutional neural network algorithm to obtain the detection target feature. The detection target feature can represent the semantic information of the data collected by the vehicle-side sensor, and thereby, the lane attribute in the detection target road area can be represented based on the detection target feature.

[0071] According to an embodiment of the present disclosure, the detection target feature represents the detection target road area based on the data collection viewing angle of the vehicle-side sensor, and the detection target feature and the target presentation information can be fused using an encoder. Based on a predetermined bird's-eye view (Bird’s Eye View, BEV, or also called a bird's-eye viewing angle) space, a viewing angle conversion can be performed on the detection target feature of the vehicle-side sensor viewing angle, and the respective attribute semantic information of the detection target feature and the target presentation information can be fused. In this way, the obtained detection target area feature in the bird's-eye view space can represent the area lane attribute of the lane boundary line in the detection target road area from the bird's-eye viewing angle, avoiding the area lane attribute representation error caused by factors such as occlusion and light irradiation in the detection target image obtained by the vehicle-side sensor based on the vehicle-side collection viewing angle, and improving the detection accuracy of the subsequent area road map based on the detection target area feature.

[0072] According to an embodiment of the present disclosure, the detection target area feature may be generated based on a plurality of detection target images detected by vehicle-side sensors of a plurality of vehicles. By identifying detection target features corresponding to each of the plurality of vehicle-side sensors, the plurality of detection target features can represent the detection target road area from the respective data collection perspectives of the plurality of vehicles. Further, the detection target area feature can represent the area lane attribute of the detection target road area based on the semantic attributes of the plurality of collection perspectives, improve the representation accuracy of the detection target road area by the detection target area feature in the bird's-eye view space, and further improve the detection accuracy of the area road map.

[0073] According to an embodiment of the present disclosure, based on an attention mechanism, a decoder is used to process the target presentation information and the detection target area feature, and based on the target presentation information, the decoder of the large-scale model can be assisted in understanding the semantic information representing the area lane attribute in the detection target area feature. Thereby, the attribute collision between the area lane attribute represented by the area road map and the related area lane attribute is reduced, and the detection accuracy and map generation efficiency of the area road map are improved.

[0074] According to an embodiment of the present disclosure, a decoder is used to process the detection target area feature and part or all of the target presentation information to obtain an area road map.

[0075] According to an embodiment of the present disclosure, the detection target feature is a multimodal detection target feature. The multimodal detection target feature is obtained by performing feature extraction on multimodal detection target information. The multimodal detection target information is detected by vehicle-side sensors located in the detection target road area, and the multimodal detection target information includes detection target images.

[0076] According to an embodiment of the present disclosure, the multimodal detection target information may include a detection target image and a detection target point cloud. The detection target point cloud may include modalities such as a lidar point cloud and a millimeter-wave radar point cloud. The detection target point cloud may be collected by vehicle-side sensors such as a lidar and a millimeter-wave radar of a vehicle. There can be a spatial correlation relationship between detection target information of different modalities, whereby alignment of the multimodal detection target information can be realized, and alignment of a plurality of multimodal detection target information obtained from a plurality of vehicles can also be realized.

[0077] According to an embodiment of the present disclosure, feature extraction can be performed on the multimodal detection target information based on any type of deep learning algorithm such as a convolutional neural network and a feature pyramid network, and the embodiments of the present disclosure do not limit this.

[0078] According to an embodiment of the present disclosure, using an encoder to fuse the detection target feature and the target presentation information to obtain a detection target area feature in the bird's-eye view space may further include using an encoder to fuse the multimodal detection target feature and the target presentation information to obtain a detection target area feature. In this way, the multimodal detection target feature is mapped to the bird's-eye view space so that the relevant lane attributes of the relevant road area represented by the target presentation information can be fused with the detection target area feature in the bird's-eye view space, and the representation accuracy of the area lanes in the detection target road area by the detection target area feature can be improved.

[0079] According to an embodiment of the present disclosure, the large-scale model may further include a detection head. The detection head may be constructed based on a neural network algorithm. For example, the detection head can be constructed based on an attention network algorithm, a convolutional neural network algorithm, or a fully connected layer.

[0080] Note that, for the purpose of explaining the data processing process of the large-scale model, the embodiments of the present disclosure can divide the large-scale model into an encoder, a decoder, and a detection head, and are not intended to limit the boundary of the specific network structure of the large-scale model. It can be understood that the detection head of the large-scale model may be included in the decoder.

[0081] According to the embodiments of the present disclosure, based on the attention mechanism, using the decoder to process the target presentation information and the detection target area features may include, based on the attention mechanism, using the decoder to fuse the lane attribute query features, the target presentation information, and the detection target area features to obtain target fusion features, and using the detection head to process the target fusion features to obtain an area road map.

[0082] According to the embodiments of the present disclosure, the lane attribute query features may be obtained by fine-tuning the large-scale model. For example, in the process of fine-tuning the model parameters of the large-scale model, the tensor of the initialized lane attribute query features can also be adjusted to obtain the lane attribute query features.

[0083] According to the embodiments of the present disclosure, taking the lane attribute query features as query features (Query), the key features (Key) and value features (Value) can be specified based on the target presentation information and the detection target area features. Furthermore, based on the cross-attention mechanism, by using the decoder pair to process the query features (Query), key features (Key), and value features (Value), the semantic information among the target presentation information, the detection target area features, and the lane attribute query features can be fully fused. Moreover, under the condition of the strong lane attribute understanding ability and expression ability of the large-scale model, the obtained target fusion features can relatively accurately represent the area lane attributes in the detection target road area, improving the detection accuracy of the area road map.

[0084] FIG. 3 schematically shows a principle diagram of a map construction method based on a large-scale model according to an embodiment of the present disclosure.

[0085] As shown in FIG. 3, the detection target road area Q300 may include a plurality of vehicles traveling in the area lanes, and a vehicle-side sensor may be attached to each of the plurality of vehicles. Each vehicle-side sensor of the plurality of vehicles can collect multi-modal detection target information such as a detection target image and a detection target point cloud. By obtaining the information respectively collected by the vehicle-side sensors of each of the plurality of vehicles in the detection target road area Q300, a plurality of detection target information P301 for input to the large-scale model can be obtained. Based on the position of the detection target road area Q300, a plurality of related area lane attributes A301 that satisfy a predetermined similarity condition with the detection target road area Q300 can also be specified.

[0086] As shown in FIG. 3, by inputting a plurality of related area lane attributes A301 and a plurality of detection target information P301 into the feature extraction network, target presentation information and a plurality of detection target features can be obtained. The plurality of detection target features and the target presentation information are input to the encoder 310 of the large-scale model, and a detection target area feature F301 in the bird's-eye view space can be output.

[0087] As shown in FIG. 3, the detection target road area Q300 may be divided into a plurality of grid sub-areas arranged in an array. For the first detection target information P3011 and the second detection target information P3012 representing the target grid sub-areas among the plurality of grid sub-areas, the encoder 310 can map the detection target information collected from the vehicle-side view angle to the bird's-eye view space to obtain a detection target area sub-feature F3011 in the bird's-eye view space. After obtaining the detection target area sub-features in the bird's-eye view space corresponding to each of the plurality of grids, it can be understood that the detection target area feature F301 in the bird's-eye view space can be specified.

[0088] Note that the first detection target information P3011 and the second detection target information P3012 may be collected by respective vehicle-side sensors of different vehicles in the detection target road area Q300, and can represent area lanes in the detection target road area from different vehicle-side viewpoints. Thus, detection target features corresponding to a plurality of vehicle-side viewpoints can be fused and mapped to an aerial view space to obtain detection target area features, realizing relatively accurate representation of the area lane attributes of the detection target road area.

[0089] As shown in FIG. 3, the detection target area feature F301 and the target presentation information are input to the decoder 320 of the large-scale model, and a target fusion feature can be output. The target fusion feature is input to the detection head 330 of the large-scale model, and an area road map G301 including any area lane attributes such as area lane boundary lines, area lane boundary line topological relationships, and area lane boundary line types can be output.

[0090] As shown in FIG. 3, the detection target area feature F301 output by the encoder 310 may be used to update the area feature library K310. The area feature library K310 can store detection target area features representing the detection target road area Q300 generated in a plurality of historical time zones.

[0091] As shown in FIG. 3, the detection target area features generated in a plurality of time zones stored in the area feature library K310 can be fused to obtain the current detection target area feature F301. Further, based on the fused detection target area feature F301, the detection target road area Q300 can be detected relatively accurately, and the accuracy of the area road map can be improved.

[0092] According to an embodiment of the present disclosure, the target presentation information may include related lane spatial features, related lane position and type features. The related lane spatial features are obtained by performing spatial feature extraction on the related area lane attributes. The related lane position and type features are determined based on the related area lane position and the related area lane type. The related area lane attributes include the related area lane position and the related area lane type.

[0093] According to an embodiment of the present disclosure, using an encoder to fuse the detection target features and the target presentation information to obtain the detection target area features may include using the first attention network of the encoder to fuse the detection target features and the related lane spatial features to obtain spatial fusion features in the bird's-eye view space, and using the second attention network of the encoder to fuse the spatial fusion features with the related lane position and type features to obtain the detection target area features.

[0094] According to an embodiment of the present disclosure, the first attention network and the second attention network may be constructed based on the Cross-attention network algorithm. By fusing the detection target features and the related lane spatial features using the first network, it is possible to relatively accurately learn the semantic information of the area lane attributes in the detection target road area in the bird's-eye view space based on the cross-attention mechanism. Then, based on the second attention network, the spatial fusion features are fused with the related lane position and type features, and the spatial fusion features in the bird's-eye view space assist in learning the area lane attributes of the area lanes in the detection target road area through the position and type of the lanes in the related road area, so that the detection target area features can relatively accurately represent important lane semantic attributes such as the type, position, and topological relationship of the area lanes, and further realize the improvement of the detection accuracy and accuracy of the subsequent area lane map.

[0095] FIG. 4 schematically shows a principle diagram of an encoder according to an embodiment of the present disclosure.

[0096] As shown in FIG. 4, the first attention network of the encoder 410 may include a first self-attention network layer 411 and a first cross-attention network layer 412, and the second attention network of the encoder 410 may include a second cross-attention network layer 421. The related area lane attributes can be represented based on the standard resolution map A401 of the related road area. The standard resolution map A401 is input into the spatial feature extraction network 401, and the related lane spatial feature F401 is output. The spatial feature extraction network 401 may be constructed based on a convolutional neural network algorithm. The standard resolution map A401 is input into the feature encoding network 402, and the related lane position and type feature F402 are output. The feature encoding network 402 can be constructed based on the position encoding layer of the transformer model so that the obtained related lane position and type feature F402 can be represented as a token sequence. It should be understood that the target presentation information may include the related lane spatial feature F401 and the related lane position and type feature F402.

[0097] As shown in FIG. 4, the related lane spatial feature F401 and the encoder query feature F411 are fused and input into the first self-attention network layer 411 to realize self-attention enhancement for the related lane spatial feature and obtain the enhanced related lane spatial feature. The enhanced related lane spatial feature is used as the query feature (key), and the detection target feature F403 is used as the value feature and the key feature and input into the first cross-attention network layer 412 to output a spatial fusion feature. The spatial fusion feature is used as the query feature, and the related lane position and type feature F402 are used as the value feature and the key feature and input into the second cross-attention network layer 421, and the detection target area feature F410' can be output.

[0098] Note that the encoder query feature F411 may be obtained by fine-tuning a large-scale model.

[0099] In one example, the relevant lane space features and the detection target area features among the target presentation information can be input into a decoder to output the target fusion features.

[0100] According to an embodiment of the present disclosure, processing the target presentation information and the detection target area features using a decoder based on an attention mechanism may include processing the lane attribute query features, the key features, and the value features using a decoder based on the attention mechanism.

[0101] According to an embodiment of the present disclosure, the key features and the value features are specified based on the detection target area features and the relevant lane space features. For example, the detection target area features and the relevant lane space features can be fused to obtain the detection target area fusion features, and the detection target area fusion features can be used as the key features and the value features.

[0102] According to an embodiment of the present disclosure, the lane attribute query features are obtained by fine-tuning a large-scale model.

[0103] According to an embodiment of the present disclosure, the detection head may include a lane boundary detection head.

[0104] According to an embodiment of the present disclosure, processing the target fusion features using the detection head to obtain the area road map may include processing the target fusion features using the lane boundary detection head to obtain the area lane boundary.

[0105] According to an embodiment of the present disclosure, the area road map is specified based on the area lane detection result, and the area lane detection result includes the area lane boundary.

[0106] According to an embodiment of the present disclosure, the area lane boundary may include the trajectory of the lane boundary in the area lane, and may further include vehicle passing rules such as the passing direction indicated by the lane boundary and the vehicle interruption method. The area lane boundary can be represented based on a vector.

[0107] According to an embodiment of the present disclosure, the detection head may further include a lane topology relationship detection head.

[0108] According to an embodiment of the present disclosure, processing the target fusion feature using the detection head to obtain the area road map may include processing the area lane boundary line and the target presentation information using the lane topology relationship detection head to obtain the area lane topology relationship.

[0109] According to an embodiment of the present disclosure, the area lane detection result further includes an area lane topology relationship. By processing the lane boundary line detection result and the target presentation information using the lane topology relationship detection head, the detection of the area lane topology relationship between the lane boundary lines in the detection target road area is assisted using the relevant lane attributes of the relevant road area represented by the target presentation information, and the detection accuracy of the area lane topology relationship can be improved.

[0110] According to an embodiment of the present disclosure, the detection head may further include a lane group detection head.

[0111] According to an embodiment of the present disclosure, processing the target fusion feature using the detection head to obtain the area road map may further include processing the target fusion feature using the lane group detection head to obtain the area lane group.

[0112] According to an embodiment of the present disclosure, the area road map is specified based on the area lane detection result, and the area lane detection result includes an area lane group representing the relevant relationship between at least two lanes in the detection target road area. For example, it may represent two adjacent area lanes having the same traffic direction.

[0113] According to an embodiment of the present disclosure, the detection head may further include a lane difference detection head.

[0114] According to an embodiment of the present disclosure, processing the target fusion feature using the detection head to obtain the area road map may further include processing the target fusion feature using the lane difference detection head to obtain the target difference information.

[0115] According to an embodiment of the present disclosure, the target difference information represents that the degree of difference between the related area lane attribute and the area lane attribute of the detected road area meets a predetermined difference condition, and the area lane detection result further includes the target difference information.

[0116] In one example, since the target difference information can represent that the difference between the area lane attribute in the currently detected road area and the related area lane attribute representing the related road area is too large, it can be estimated that due to reasons such as road construction and changes in road driving rules, it may be difficult for the related area lane attribute to accurately represent the current related road area. In this way, using the detection target information (including at least one of the detection target image and the detection target point cloud) collected by the vehicle-side sensor, the related area lane attribute that needs to be updated can be identified relatively accurately. Furthermore, using at least one of the detection target image and the detection target point cloud related to the related road area, the related area road map of the related road area can be generated, realizing accurate update of map information, and improving the accuracy and safety of the vehicle executing the automatic driving function in a plurality of related road areas.

[0117] According to an embodiment of the present disclosure, the map construction method based on the large-scale model may further include updating at least one related area lane attribute of the related area lane boundary line, the related area lane topology relationship, and the related area lane group based on the area lane detection result when the target difference information is obtained.

[0118] In one example, based on the area lane boundary line, attributes of the lane boundary line such as the lane position and lane type of the related area lane boundary line can be updated, and further rapid calibration and update of the related area lane boundary line can be realized.

[0119] In one example, the related area lane topology relationship can be updated based on the area lane topology relationship.

[0120] In one example, the related area lane group may be updated based on the area lane group.

[0121] It should be understood that, based on at least one of the area lane boundary line, the area lane topology relationship, and the area lane group, the corresponding related area lane attribute can be updated. In this way, by using the detection target information obtained by the highly real-time vehicle-side sensor, the related area lane attribute of the related road area is updated in real time, the lane attribute update efficiency of the entire road passage range including the detection target road area and the related road area is improved, and the traffic safety and traffic efficiency of vehicles in the entire road passage range can be improved.

[0122] FIG. 5 schematically shows a principle diagram of a map construction method based on a large-scale model according to another embodiment of the present disclosure.

[0123] As shown in FIG. 5, the decoder 520 of the large-scale model may include a decoder self-attention network 521 and a decoder attention network 522. The detection head may include a lane boundary line detection head 531, a lane topology relationship detection head 532, a lane group detection head 533, a lane difference detection head 534, and a lane boundary line segmentation detection head 535. The decoder attention network 522 may be constructed based on a cross-attention network algorithm.

[0124] As shown in FIG. 5, the lane attribute query feature F511 obtained after fine-tuning is input into the decoder self-attention network 521 to obtain an enhanced query feature. After fusing the detection target area feature F501 and the target presentation information F502, key features and value features are generated. The enhanced query feature, key feature, and value feature are input into the decoder attention network 522 to obtain a target fusion feature F521.

[0125] As shown in FIG. 5, the target fusion feature F521 is input to the lane boundary detection head 531 to obtain the area lane boundary (polyline). The area lane boundary and the target presentation information F502 are input to the lane topology relationship detection head 532, and the area lane topology relationship can be obtained. The target fusion feature F521 is input to the lane group detection head 533 to obtain the area lane group. The target fusion feature F521 is input to the lane difference detection head 534, and the target difference information can be obtained. The target fusion feature F521 and the detection target area feature 501 are input to the lane boundary segmentation detection head 535, and the segmentation area between the foreground and the background in the detection target image can be output, and further the area lane boundary can be represented based on the segmentation area.

[0126] Note that the detection target area feature 501 and the target presentation information 502 surrounded by the dotted line frame shown in FIG. 5 may be the same as the detection target area feature 501 and the target presentation information 502 surrounded by the solid line frame.

[0127] In one example, the lane attribute query feature may include three query feature vectors: a lane boundary query feature, a lane group query feature, and a lane boundary segmentation query feature. Different query feature vectors are input to different decoder self-attention networks, and attention fusion is performed with the detection target area feature and the target presentation information respectively using the three enhanced query feature vectors, and three different initial fusion features can be obtained. By overlapping different initial fusion features, the target fusion feature is obtained.

[0128] According to an embodiment of the present disclosure, the map construction method based on a large-scale model may further include constructing a target navigation map based on each area road map of a plurality of detection target road areas.

[0129] According to an embodiment of the present disclosure, a plurality of roads to be detected may belong to the same driving range. For example, they may belong to the same working area range or the driving range of the same city. By aggregating the area road maps of each of the plurality of roads to be detected, a target navigation map representing the overall driving range is constructed, and it is realized to assist an autonomous vehicle in executing an autonomous driving function using the target navigation map, thereby improving the driving efficiency and safety of the vehicle.

[0130] Based on the above map construction method based on a large-scale model, an embodiment of the present disclosure further provides a vehicle control method. Hereinafter, it will be described with reference to the drawings and embodiments.

[0131] It should be noted that the processes of obtaining, collecting, or processing information according to the embodiments of the present disclosure are all executed on the condition of obtaining the permission of relevant users or institutions, and relevant users or institutions are clearly notified that the purpose of the method provided by the embodiments of the present disclosure is to improve the driving safety and driving efficiency of vehicles. The method provided by the embodiments of the present disclosure adopts necessary encryption or desensitization measures for the obtained information to avoid information leakage. The obtained information includes, but is not limited to, information to be detected, lane attributes of related areas, etc.

[0132] FIG. 6 schematically shows a flowchart of a vehicle control method according to an embodiment of the present disclosure.

[0133] As shown in FIG. 6, the vehicle control method may include operation S610.

[0134] In operation S610, the driving of the vehicle is controlled based on the road map.

[0135] According to an embodiment of the present disclosure, a road map may be constructed by a map construction method based on a large-scale model provided by the embodiment of the present disclosure. For example, the road map may be an area road map obtained by a map construction method based on a large-scale model provided by the embodiment of the present disclosure. Further, for example, the road map may be a target navigation map obtained by a map construction method based on a large-scale model provided by the embodiment of the present disclosure.

[0136] According to an embodiment of the present disclosure, by obtaining a road map identified by a map construction method based on a large-scale model provided by the embodiment of the present disclosure, a vehicle can complete a lane-level automatic driving function with the support of a relatively accurate and timely road map, and the safety and driving efficiency of vehicle driving can be improved.

[0137] FIG. 7 schematically shows a block diagram of a map construction apparatus based on a large-scale model according to an embodiment of the present disclosure.

[0138] As shown in FIG. 7, the map construction apparatus 700 based on a large-scale model may include an acquisition module 710, a target presentation information construction module 720, and a first map construction module 730.

[0139] The acquisition module 710 is used to acquire relevant area lane attributes and a detection target image collected by a vehicle-side sensor. Here, the detection target image represents a detection target road area, the relevant area lane attributes correspond to a relevant road area, and a predetermined similarity condition is satisfied between the relevant road area and the detection target road area.

[0140] The target presentation information construction module 720 is used to construct target presentation information based on the relevant area lane attributes.

[0141] The first map construction module 730 is used to process the target presentation information and the detection target image using a large-scale model to obtain an area road map of the detection target road area.

[0142] According to an embodiment of the present disclosure, the large-scale model includes an encoder and a decoder.

[0143] According to an embodiment of the present disclosure, the first map construction module includes a detection target area feature acquisition sub-module and an area road map acquisition sub-module.

[0144] The detection target area feature acquisition sub-module is used to fuse the detection target feature and the target presentation information using an encoder to obtain the detection target area feature in the bird's-eye view space, where the detection target feature is specified based on the detection target image.

[0145] The area road map acquisition sub-module processes the target presentation information and the detection target area feature using a decoder based on an attention mechanism to obtain an area road map.

[0146] According to an embodiment of the present disclosure, the large-scale model further includes a detection head.

[0147] According to an embodiment of the present disclosure, the area road map acquisition sub-module includes a target fusion feature acquisition means and an area road map acquisition means.

[0148] The target fusion feature acquisition means is used to fuse the lane attribute query feature, the target presentation information, and the detection target area feature using a decoder based on an attention mechanism to obtain a target fusion feature, where the lane attribute query feature is obtained by fine-tuning the large-scale model.

[0149] The area road map acquisition means processes the target fusion feature using a detection head to obtain an area road map.

[0150] According to an embodiment of the present disclosure, the detection head includes a lane boundary detection head.

[0151] According to an embodiment of the present disclosure, the area road map acquisition means includes an area lane boundary acquisition sub-means.

[0152] The area lane boundary line acquisition sub-means is used to process the target fusion feature using a lane boundary line detection head to obtain an area lane boundary line. Here, the area road map is specified based on the area lane detection result, and the area lane detection result includes the area lane boundary line.

[0153] According to an embodiment of the present disclosure, the detection head further includes a lane topology relationship detection head.

[0154] According to an embodiment of the present disclosure, the area road map acquisition means further includes an area lane topology relationship acquisition sub-means.

[0155] The area lane topology relationship acquisition sub-means is used to process the area lane boundary line and the target presentation information using a lane topology relationship detection head to obtain an area lane topology relationship. Here, the area lane detection result further includes the area lane topology relationship.

[0156] According to an embodiment of the present disclosure, the detection head includes a lane group detection head.

[0157] According to an embodiment of the present disclosure, the area road map acquisition means includes an area lane group acquisition sub-means.

[0158] The area lane group acquisition sub-means is used to process the target fusion feature using a lane group detection head to obtain an area lane group. Here, the area road map is specified based on the area lane detection result, and the area lane detection result includes an area lane group representing the relationship between at least two lanes in the detection target road area.

[0159] According to an embodiment of the present disclosure, the detection head includes a lane difference detection head.

[0160] According to an embodiment of the present disclosure, the area road map acquisition means includes a target difference information acquisition sub-means.

[0161] The target difference information acquisition sub-means is used to process the target fusion features using a lane difference detection head to obtain target difference information. Here, the target difference information indicates that the degree of difference between the related area lane attribute and the area lane attribute of the detection target road area satisfies a predetermined difference condition, and the area lane detection result further includes the target difference information.

[0162] According to an embodiment of the present disclosure, the map construction device 700 based on a large-scale model further includes an update module.

[0163] When the update module obtains the target difference information, it is used to update at least one related area lane attribute among the related area lane boundary line, the related area lane topology relationship, and the related area lane group based on the area lane detection result.

[0164] According to an embodiment of the present disclosure, the target presentation information includes related lane spatial features and related lane position and type features. The related lane spatial features are obtained by performing spatial feature extraction on the related area lane attribute, and the related lane position and type features are specified based on the related area lane position and the related area lane type. The related area lane attribute includes the related area lane position and the related area lane type.

[0165] According to an embodiment of the present disclosure, the detection target area feature acquisition sub-module includes a spatial fusion feature acquisition means and a detection target area feature acquisition means.

[0166] The spatial fusion feature acquisition means is used to fuse the detection target features and the related lane spatial features using the first attention network of the encoder to obtain spatial fusion features in the bird's-eye view space.

[0167] The detection target area feature acquisition means is used to fuse the spatial fusion features and the related lane position and type features using the second attention network of the encoder to obtain the detection target area features.

[0168] According to an embodiment of the present disclosure, the area road map acquisition sub-module includes a processing means.

[0169] The processing means is used to process the lane attribute query feature, key feature, and value feature using a decoder based on an attention mechanism, where the key feature and value feature are specified based on the detection target area feature and the related lane space feature, and the lane attribute query feature is obtained by fine-tuning a large-scale model.

[0170] According to an embodiment of the present disclosure, the detection target feature is a multi-modal detection target feature, the multi-modal detection target feature is obtained by performing feature extraction on multi-modal detection target information, the multi-modal detection target information is detected by a vehicle-side sensor located in the detection target road area, and the multi-modal detection target information includes a detection target image.

[0171] According to an embodiment of the present disclosure, the related area lane attribute includes a related area lane position and a related area lane type.

[0172] According to an embodiment of the present disclosure, the target presentation information construction module includes a feature fusion sub-module and a target presentation information construction sub-module.

[0173] The feature fusion sub-module is used to perform feature fusion on the related area lane position and the related area lane type to obtain the related lane position and type features.

[0174] The target presentation information construction sub-module is used to construct target presentation information based on the related lane position and type features.

[0175] According to an embodiment of the present disclosure, a plurality of related road areas are included, and the related area lane attributes of the plurality of related road areas correspond to a plurality of related lane positions and type features.

[0176] According to an embodiment of the present disclosure, the target presentation information construction sub-module includes a relevance weight determination means and a target presentation information acquisition means.

[0177] The relevance weight determination means is used to determine the respective relevance weights of a plurality of relevant lane positions and type features based on at least one relevant lane position and type feature, and the relevance weight represents the degree of relevance between the relevant road area and the detection target road area.

[0178] The target presentation information acquisition means fuses a plurality of relevant lane positions and type features based on a plurality of relevance weights to acquire target presentation information.

[0179] According to an embodiment of the present disclosure, the map construction device 700 based on a large-scale model further includes a second map construction module.

[0180] The second map construction module is used to construct a target navigation map based on the area road map of each of a plurality of detection target road areas.

[0181] FIG. 8 schematically shows a block diagram of a vehicle control device according to an embodiment of the present disclosure.

[0182] As shown in FIG. 8, the vehicle control device 800 includes a vehicle control module 810.

[0183] The vehicle control module 810 is used to control the running of the vehicle based on a road map, where the road map is constructed by a map construction device based on a large-scale model provided by an embodiment of the present disclosure.

[0184] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.

[0185] According to an embodiment of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a map construction method based on a large-scale model provided by the embodiment of the present disclosure.

[0186] According to an embodiment of the present disclosure, there is provided an electronic device including at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a vehicle control method provided by the embodiment of the present disclosure.

[0187] According to an embodiment of the present disclosure, there is provided an autonomous vehicle including an electronic device for executing a vehicle control method provided by the embodiment of the present disclosure.

[0188] According to an embodiment of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute a method provided by the embodiment of the present disclosure.

[0189] According to an embodiment of the present disclosure, there is provided a computer program that, when executed by a processor, implements a method provided by the embodiment of the present disclosure.

[0190] FIG. 9 schematically shows a block diagram of an electronic device suitable for implementing a map construction method and a vehicle control method based on a large model according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may further represent various forms of mobile devices, for example, personal digital assistants, mobile phones, smartphones, wearable devices, and other similar computing devices. The members, their connections and relationships, and their functions shown in this specification are merely illustrative and do not limit the implementation of the present disclosure described and / or claimed in this specification.

[0191] As shown in FIG. 9, the device 900 includes a computing unit 901, and the computing unit 901 may execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may further store various programs and data necessary for the operation of the device 900. The computing unit 901, the ROM 902, and the RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0192] A plurality of components in the device 900 are connected to the I / O interface 905, including an input unit 906 such as a keyboard and a mouse, an output unit 907 such as various types of displays and speakers, a storage unit 908 such as a magnetic disk and an optical disk, and a communication unit 909 such as a network card, a modem, and a wireless communication transceiver. The communication unit 909 enables the device 900 to exchange information and data with other devices via a computer network such as the Internet and / or various electrical networks.

[0193] The computing unit 901 may be various general-purpose and / or dedicated processing modules having processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a GPU (Graphics Processing Unit), various dedicated artificial intelligence (AI) computing chips, computing units running various machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 901 executes each of the methods and processes described above, for example, the map construction method based on a large-scale model and the vehicle control method. For example, in some embodiments, the map construction method based on a large-scale model and the vehicle control method may be realized as a computer software program tangibly included in a machine-readable medium such as the storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed in the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the map construction method and the vehicle control method based on the large-scale model described above may be executed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute the map construction method and the vehicle control method based on a large-scale model in any other suitable manner (for example, via firmware).

[0194] The various embodiments of the systems and techniques described above in this specification may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be implemented in one or more computer programs, which may be executed and / or interpreted in a programmable system including at least one programmable processor, where the programmable processor may be a dedicated or general-purpose programmable processor, and which includes receiving data and instructions from, and transmitting data and instructions to, a memory system, at least one input device, and at least one output device.

[0195] The program code for implementing the methods of the present disclosure may be created in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions and operations defined in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the device, partially on the device, partially on the device as an independent software package, and partially on a remote device or entirely on a remote device or server.

[0196] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use in or in combination with an instruction execution system, apparatus, or electronic device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium include electrical connections by one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0197] To provide for interaction with a user, a computer may implement the systems and techniques described herein, the computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other kinds of devices may be further provided for interacting with the user, for example, feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form (including voice input, speech input, or tactile input).

[0198] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser, through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such background components, middleware components, or front-end components. The components of the system can be connected to each other by digital data communication in any form or medium (e.g., a communication network). Exemplary communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0199] The computer system may include a client and a server. The client and the server are generally separated and usually interact via a communication network. The relationship between the client and the server is generated by a computer program running on the corresponding computer and having a client-server relationship. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0200] It should be understood that various forms of the flows shown above may be used, and the operations may be sorted, added, or deleted again. For example, each operation described in the present disclosure may be executed in parallel, sequentially, or in a different order, and the present specification is not limited herein as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved.

[0201] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and alternatives can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.

Claims

1. Acquiring related area lane attributes and a detection target image collected by a vehicle-side sensor, the detection target image represents a detection target road area, the related area lane attributes correspond to the related road area, and a predetermined similarity condition is satisfied between the related road area and the detection target road area; constructing landmark presentation information based on the relevant area lane attributes; and processing the target presentation information and the target image using a large scale model to obtain an area road map of the target road area. A map building method based on large-scale models.

2. the large scale model includes an encoder and a decoder; Processing the target presentation information and the detection target image using a large scale model to obtain an area road map of the detection target road area, Using the encoder, a detection target feature and the target presentation information are combined to obtain a detection target area feature in a bird's-eye view space, and the detection target feature is identified based on the detection target image; and and processing the target presentation information and the detection target area features using the decoder based on an attention mechanism to obtain the area road map. The method of claim 1.

3. The large scale model further includes a detector head; processing the target presentation information and the detection target area features using the decoder based on an attention mechanism, Using the decoder to fuse lane attribute query features, the target presentation information, and the detection target area features based on an attention mechanism to obtain target fusion features, where the lane attribute query features are obtained by fine-tuning the large-scale model; and processing the target fusion features with the detection head to obtain the area road map. The method of claim 2.

4. The detection head includes a lane boundary detection head, Processing the target fusion features with the detection head to obtain the area road map, and processing the target fusion features using the lane boundary detection head to obtain area lane boundaries, the area road map being identified based on the area lane detection result, the area lane detection result including the area lane boundaries. The method according to claim 3.

5. The detection head further includes a lane topology relationship detection head; Processing the target fusion features with the detection head to obtain the area road map, The method further includes using the lane topology relationship detection head to process the area lane boundary line and the target presentation information to obtain an area lane topology relationship, and the area lane detection result further includes the area lane topology relationship; The method according to claim 4.

6. The detection head includes a lane group detection head; Processing the target fusion features with the detection head to obtain the area road map, and processing the target fusion feature using the lane group detection head to obtain an area lane group, the area road map being identified based on an area lane detection result, the area lane detection result including the area lane group, the area lane group representing an association relationship between at least two lanes in the detected road area. The method according to claim 3.

7. The detection head includes a lane difference detection head, Processing the target fusion features with the detection head to obtain the area road map, and processing the target fusion feature using the lane difference detection head to obtain target difference information, the target difference information indicating that a difference degree between the related area lane attribute and the area lane attribute of the detection target road area satisfies a predetermined difference condition, and the area lane detection result further includes the target difference information. The method according to claim 4.

8. When the target difference information is obtained, the method further includes updating at least one related area lane attribute of a related area lane boundary line, a related area lane topological relationship, and a related area lane group according to the area lane detection result. The method of claim 7.

9. The target presentation information includes related lane spatial features and related lane position and type features, the related lane spatial features being obtained by performing spatial feature extraction on the related area lane attributes, the related lane position and type features being identified based on the related area lane position and the related area lane type, the related area lane attributes including the related area lane position and the related area lane type, Fusing the detection target feature and the target presentation information using the encoder to obtain a detection target area feature, Fusing the detection target feature and the associated lane spatial feature using a first attention network of the encoder to obtain a spatial fusion feature in a bird's-eye view space; and fusing the spatial fusion feature with the associated lane position and type features using a second attention network of the encoder to obtain the detected target area feature. The method of claim 2.

10. processing the target presentation information and the detection target area features using the decoder based on the attention mechanism, and processing lane attribute query features, key features and value features using the decoder based on an attention mechanism, the key features and the value features being identified based on the detection target area features and the associated lane spatial features, and the lane attribute query features being obtained by fine-tuning the large-scale model.

10. The method of claim 9.

11. the detection target features are multimodal detection target features, the multimodal detection target features are obtained by performing feature extraction on multimodal detection target information, the multimodal detection target information is detected by a vehicle-side sensor located in the detection target road area, and the multimodal detection target information includes the detection target image; The method of claim 2.

12. The related area lane attributes include a related area lane position and a related area lane type; Constructing the target presentation information based on the relevant area lane attributes includes: performing feature fusion on the relevant area lane position and the relevant area lane type to obtain relevant lane position and type features; and constructing the target presentation information based on the associated lane position and type characteristics. The method of claim 1.

13. A plurality of related road areas are included, and related area lane attributes of the plurality of related road areas correspond to a plurality of the related lane position and type features; Constructing the target presentation information based on the associated lane position and type characteristics includes: determining a relevance weight for each of the plurality of relevant lane position and type features based on at least one of the relevant lane position and type features, the relevance weight representing a degree of relevance between the relevant road area and the detected road area; and fusing the plurality of relevant lane position and type features based on the plurality of relevance weights to obtain the target presentation information. The method of claim 12.

14. and constructing a target navigation map based on the area road map of each of the plurality of detected road areas. The method of claim 1.

15. and controlling vehicle travel based on a road map constructed by the method according to any one of claims 1 to 14. A vehicle control method.

16. an acquisition module for acquiring related area lane attributes and a detection target image collected by a vehicle-side sensor, the detection target image representing a detection target road area, the related area lane attributes corresponding to a related road area, and a predetermined similarity condition being satisfied between the related road area and the detection target road area; A target presentation information construction module constructs target presentation information based on the related area lane attributes; a first map construction module for processing the target presentation information and the target image using a large scale model to obtain an area road map of the target road area; A map-building device based on large-scale models.

17. the large scale model includes an encoder and a decoder; The first map construction module includes: a detection target area feature acquisition submodule that uses the encoder to fuse a detection target feature and the target presentation information to acquire a detection target area feature in a bird's-eye view space, the detection target feature being identified based on the detection target image; and an area road map acquisition submodule for processing the target presentation information and the detection target area features using the decoder based on an attention mechanism to acquire the area road map.

17. The apparatus of claim 16.

18. The large scale model further includes a detector head; The area road map acquisition submodule: a target fusion feature acquisition means for acquiring a target fusion feature by fusing the lane attribute query feature, the target presentation information and the detection target area feature using the decoder based on an attention mechanism, the lane attribute query feature being obtained by fine-tuning the large-scale model; and an area road map acquisition means for processing the target fusion features using the detection head to acquire the area road map.

20. The apparatus of claim 17.

19. The detection head includes a lane boundary detection head, The area road map acquisition means and an area lane boundary line acquisition sub-means for processing the target fusion feature using the lane boundary line detection head to acquire area lane boundary lines, the area road map being identified based on the area lane detection result, and the area lane detection result including the area lane boundary lines.

20. The apparatus of claim 18.

20. The detection head further includes a lane topology relationship detection head; The area road map acquisition means Further includes an area lane topology relationship acquisition sub-means for processing the area lane boundary line and the target presentation information using the lane topology relationship detection head to acquire an area lane topology relationship, and the area lane detection result further includes the area lane topology relationship.

20. The apparatus of claim 19.

21. The detection head includes a lane group detection head; The area road map acquisition means and an area lane group acquisition sub-means for processing the target fusion feature using the lane group detection head to acquire an area lane group, the area road map being identified based on an area lane detection result, the area lane detection result including the area lane group, and the area lane group representing an association relationship between at least two lanes in the detection target road area.

20. The apparatus of claim 18.

22. The detection head includes a lane difference detection head, The area road map acquisition means a target difference information acquisition sub-means for processing the target fusion feature using the lane difference detection head to acquire target difference information, the target difference information representing that the difference degree between the related area lane attribute and the area lane attribute of the detection target road area satisfies a predetermined difference condition, and the area lane detection result further includes the target difference information; 20. The apparatus of claim 19.

23. Further comprising an update module for updating at least one related area lane attribute of a related area lane boundary line, a related area lane topological relationship, and a related area lane group according to the area lane detection result when the target difference information is obtained; 23. The apparatus of claim 22.

24. The target presentation information includes related lane spatial features and related lane position and type features, the related lane spatial features being obtained by performing spatial feature extraction on the related area lane attributes, the related lane position and type features being identified based on the related area lane position and the related area lane type, the related area lane attributes including the related area lane position and the related area lane type, The detection target area feature acquisition submodule: a spatial fusion feature acquisition means for acquiring a spatial fusion feature in a bird's-eye view space by fusing the detection target feature and the associated lane spatial feature using a first attention network of the encoder; and a detection target area feature acquisition means for fusing the spatial fusion feature and the associated lane position and type feature using a second attention network of the encoder to acquire the detection target area feature.

20. The apparatus of claim 17.

25. The area road map acquisition submodule: processing means for processing lane attribute query features, key features and value features using the decoder based on an attention mechanism, the key features and the value features being identified based on the detection target area features and the associated lane spatial features, and the lane attribute query features being obtained by fine-tuning the large-scale model; 25. The apparatus of claim 24.

26. the detection target features are multimodal detection target features, the multimodal detection target features are obtained by performing feature extraction on multimodal detection target information, the multimodal detection target information is detected by a vehicle-side sensor located in the detection target road area, and the multimodal detection target information includes the detection target image; 20. The apparatus of claim 17.

27. The related area lane attributes include a related area lane position and a related area lane type; The goal presentation information construction module includes: a feature fusion submodule for performing feature fusion on the relevant area lane position and the relevant area lane type to obtain relevant lane position and type features; and a target presentation information construction submodule for constructing the target presentation information based on the relevant lane position and type characteristics.

17. The apparatus of claim 16.

28. A plurality of related road areas are included, and related area lane attributes of the plurality of related road areas correspond to a plurality of the related lane position and type features; The target presentation information construction submodule includes: a relevance weight specifying means for specifying a relevance weight of each of the plurality of related lane position and type features based on at least one of the related lane position and type features, the relevance weight representing a degree of relevance between the related road area and the detection target road area; and a target presentation information acquisition means for acquiring the target presentation information by fusing the plurality of related lane position and type features based on the plurality of related lane weights.

28. The apparatus of claim 27.

29. and a second map construction module for constructing a target navigation map based on the area road map of each of the plurality of detected road areas.

17. The apparatus of claim 16.

30. A vehicle control module that controls vehicle travel based on a road map constructed by the device according to any one of claims 16 to 29. Vehicle control device.

31. At least one processor; a memory in communication with the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs the method according to any one of claims 1 to 14. electronic equipment.

32. At least one processor; a memory in communication with the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs the method of claim 15. electronic equipment.

33. The electronic device according to claim 32, Self-driving vehicles.

34. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions are used to cause a computer to carry out the method according to any one of claims 1 to 14, A non-transitory computer-readable storage medium.

35. A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 14.

36. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions are used to cause a computer to carry out the method of claim 15. A non-transitory computer-readable storage medium.

37. A computer program product which, when executed by a processor, implements the method according to claim 15.

Citation Information

Patent Citations

  • Vehicle track prediction method, device and system and storage medium

    CN116682018A

  • Map generation system, server, vehicle-side device, method, and storage medium

    JP2020038361A

  • Dynamic map rendering

    US20210199443A1