A high-precision map reconstruction method
The circumferential image features are extracted through the backbone network and combined with the bird's-eye view features and map vectors for decoding and updating, which solves the problem that traditional map construction methods are difficult to meet the requirements of high precision and real-time, and achieves high-precision and reliability map reconstruction effect.
Patent Information
- Application Number
- CN202510291138.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-03-12
AI Technical Summary
Traditional map construction methods are difficult to meet the requirements of high precision, real-time and adaptability. Deep learning technology has made progress in image processing and feature extraction, but it still faces challenges such as multi-view image information integration and uncertainty processing in map reconstruction applications.
The front view features of the surrounding view image are extracted through the backbone network and converted into bird's-eye view features. The front view and bird's-eye view features are decoded and updated with the initial map vector to generate high-precision map reconstruction results.
It significantly improves the accuracy and reliability of map reconstruction, effectively reduces information losses, deals with uncertain factors, adapts to complex and changeable actual environments, and meets high-precision and real-time requirements.
Smart Images

Figure CN119784960B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of autonomous driving, and specifically provides a high-precision map reconstruction method. Background Art
[0002] With the rapid development of fields such as autonomous driving technology and intelligent transportation systems, the demand for high-precision maps as key infrastructure is increasing day by day. Traditional map construction methods are difficult to meet the requirements of high precision, real-time performance, and adaptability. Although deep learning technology has made remarkable progress in image processing and feature extraction, it still faces many challenges in map reconstruction applications, such as how to effectively integrate multi-view image information, accurately handle uncertain factors, and improve the accuracy and reliability of the reconstructed map. Summary of the Invention
[0003] The purpose of the present invention is to provide a high-precision map reconstruction method to solve the problems raised in the above background art.
[0004] To achieve the above purpose, the present invention provides the following technical solutions:
[0005] In a first aspect, the present invention provides a high-precision map reconstruction method, including the following steps:
[0006] Extract the front view features of the surround-view image through the backbone network;
[0007] Convert the front view features into bird's-eye view features;
[0008] Decode the front view features and the initialized map vector to obtain two-dimensional front view object-level information and point-level information, and use them as formal map prompts;
[0009] Combine the formal map prompts, the initialized map vector, and the bird's-eye view features to update the map vector and the bird's-eye view features;
[0010] Obtain the updated map vector and bird's-eye view features to generate a high-precision map reconstruction result.
[0011] Preferably, the surround-view image refers to an image sequence of the surrounding environment collected by a surround-view camera installed on a vehicle or other mobile platform, and is obtained by preprocessing the image sequence.
[0012] Preferably, the preprocessing of the image sequence includes image denoising, grayscale conversion, cropping, and scaling operations to improve the image quality and make it suitable for input into the backbone network.
[0013] Preferably, the extraction of the front view features of the surround view image through the backbone network means inputting the preprocessed surround view image into the backbone network, and the backbone network extracts features layer by layer from the image according to its pre-trained model parameters and convolutional neural network structure, and finally outputs the front view features.
[0014] Preferably, obtaining the two-dimensional front view object-level information and point-level information from the input front view features and the initialized map vector, and using them as the formal map prompt words includes:
[0015] Initializing the map vector and decoding it together with the front view features;
[0016] Through decoding calculations, obtain the information of the initial reference point and uncertainty of the front view image, and this information constitutes a part of the front view prompt words;
[0017] At the same time, convert the front view features into bird's-eye view features;
[0018] Use the front view prompt words, the initialized map vector, and the bird's-eye view features as two-dimensional uncertainty prompt words to generate an updated map vector and bird's-eye view features, further enriching the prompt word information.
[0019] In a second aspect, the present invention provides a high-precision map reconstruction system for implementing the method of any of the above embodiments, and the system includes:
[0020] A backbone network for extracting the front view features of the surround view image;
[0021] A front view to bird's-eye view module for converting the front view features into bird's-eye view features;
[0022] A front view uncertainty-aware decoder for decoding the front view features and the initialized map vector to obtain two-dimensional front view object-level information and point-level information, and using them as formal map prompt words;
[0023] A two-dimensional uncertainty prompt word module for combining the formal map prompt words, the initialized map vector, and the bird's-eye view features to update the map vector and the bird's-eye view features;
[0024] A bird's-eye view uncertainty-aware decoder for generating a high-precision map reconstruction result after obtaining the updated map vector and bird's-eye view features.
[0025] In a third aspect, the present invention provides an electronic device, including at least one processor, and the processor is communicatively connected to at least one memory. Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any of the above embodiments.
[0026] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer instructions for causing a processor to implement the method described in any of the above embodiments when executed.
[0027] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0028] 1. Improve map accuracy: Through the collaborative work of the backbone network, the PV2BEV module, and two UA-decoders, the multi-view information of the surround-view images can be fully utilized to extract and integrate map information from different angles, thereby significantly improving the accuracy of the reconstructed map.
[0029] 2. The UA-attention module in the UA-decoder and the cross-attention modules (P2BEV-attention and P2Q-attention) in the UI2DPrompt can focus on key features, effectively reduce information loss, and further improve the accuracy of map reconstruction.
[0030] 3. Handle uncertainty: An uncertainty perception mechanism is introduced to quantify and process uncertain factors during map reconstruction; by calculating the uncertainty of the front-view image and integrating it into the prompt generation and map reconstruction processes, the reliability of the map reconstruction result can be improved, making it better adapt to complex and changing actual environments.
[0031] 4. High efficiency and adaptability: The method of the present invention is based on a deep learning model and can automatically learn and adapt to the map reconstruction requirements in different scenarios. The design of each module has a certain degree of generality and flexibility, and can be adjusted and optimized according to the actual application scenario, improving the adaptability and practicality of the method.
[0032] 5. Adopting a modular design enables each module to perform parallel computing or targeted optimization, improving the efficiency of the entire map reconstruction process and meeting application scenarios with high real-time requirements, such as real-time map updates for autonomous driving vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a flowchart of a high-precision map reconstruction method according to an embodiment of the present invention;
[0034] Figure 2 is a sub - flowchart of step S300;
[0035] Figure 3 is the training and inference flowchart of a high - precision map reconstruction system according to an embodiment of the present invention;
[0036] Figure 4 is the working principle diagram of the front - view uncertainty encoder of the present invention;
[0037] Figure 5 is the working principle diagram of the two - dimensional uncertainty prompt word module 400 of the present invention;
[0038] Figure 6 is the working principle diagram of the front - view uncertainty perception decoder 300 of the present invention;
[0039] Figure 7 is the module diagram of a high - precision map reconstruction system according to an embodiment of the present invention;
[0040] Figure 8 is the module diagram of the electronic device according to an embodiment of the present invention. Detailed implementation manners
[0041] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] Please refer to Figure 1 , the present invention provides a technical solution: a high - precision map reconstruction method, including the following steps:
[0043] S100: Extract the front - view features of the surround - view image through the backbone network;
[0044] S200: Convert the front - view features into bird's - eye view features;
[0045] S300: Decode the front - view features and the initialized map vector to obtain two - dimensional front - view object - level information and point - level information, and use them as the formal map prompt words;
[0046] S400: Combine the formal map prompt words, the initialized map vector and the bird's - eye view features to update the map vector and the bird's - eye view features;
[0047] S500: Obtain the updated map vector and bird's - eye view features to generate a high - precision map reconstruction result.
[0048] In one embodiment of the present invention, the surround view image in step S100 refers to an image sequence of the surrounding environment collected by surround view cameras installed on a vehicle or other mobile platform, and is obtained by preprocessing the image sequence.
[0049] In one embodiment of the present invention, the preprocessing of the image sequence includes image denoising, grayscale conversion, cropping, and scaling operations to improve the image quality and make it suitable for input into the backbone network.
[0050] In one embodiment of the present invention, the backbone network adopts a deep convolutional neural network architecture for extracting front view features from the input surround view image. This network is pre-trained and optimized to effectively capture key information in the image and provide rich feature representations for subsequent modules.
[0051] In one embodiment of the present invention, the specific operation of step S100 refers to inputting the preprocessed surround view image into the backbone network. The backbone network extracts features layer by layer from the image according to its pre-trained model parameters and convolutional neural network structure, and finally outputs the front view features.
[0052] In one embodiment of the present invention, please refer to Figure 2 , step S300 specifically includes:
[0053] S301: Initialize the map vector and decode it together with the front view features;
[0054] S302: Through decoding calculations, obtain information on the initial reference point and uncertainty of the front view image, which constitutes part of the front view prompt;
[0055] S303: At the same time, convert the front view features into bird's-eye view features;
[0056] S304: Use the front view prompt, the initialized map vector, and the bird's-eye view features as two-dimensional uncertainty prompts to generate an updated map vector and bird's-eye view features, further enriching the prompt information.
[0057] Please refer to Figure 7 , the present application also provides a high-precision map reconstruction system, which includes:
[0058] A backbone network 100 for extracting front view features of the surround view image;
[0059] A front view to bird's-eye view module (PV2BEV) 200 for converting front view features into bird's-eye view features;
[0060] The front view uncertainty-aware decoder UA-decoder(PV) 300 decodes the front view features and the initialized map vector to obtain two-dimensional front view object-level information and point-level information, which are used as the front view prompt words;
[0061] The two-dimensional uncertainty prompt word module (UI2DPromp) 400 combines the front view prompt words, the initialized map vector, and the bird's-eye view features to update the map vector and the bird's-eye view features;
[0062] The bird's-eye view uncertainty-aware decoder UA-decoder(BEV) 500 generates a high-precision map reconstruction result after obtaining the updated map vector and the bird's-eye view features.
[0063] In an embodiment of the present invention, the front view uncertainty-aware decoder 300 includes an uncertainty-aware attention module (UA-attention); the two-dimensional uncertainty prompt word module 400 includes a hybrid injection module (Hybridinjection) 401, a front view to bird's-eye view attention module (P2BEV-attention) 402, and a front view transformation vector module (P2Q-attention) 403. Both P2BEV-attention and P2Q-attention are cross-attention modules.
[0064] In an embodiment of the present invention, please refer to Figure 3 This system is through the training and inference processes and the training-only process in the process of implementing a high-precision map reconstruction method in the above embodiment.
[0065] In an embodiment of the present invention, please refer to Figure 4 The working principle of the front view uncertainty-aware decoder 300: Input the image features extracted from the panoramic view by the backbone network and the initialized map vector into the front view uncertainty-aware decoder 300. The initialized map vector is input into the uncertainty-aware attention module after passing through the self-attention module and normalization to obtain an updated map vector. Multiply the updated map vector and the sampled image features to obtain the final map vector. Based on this vector, the initial reference point and uncertainty of the front view image can be obtained.
[0066] In an embodiment of the present invention, please refer to FIG. 5. The working principle of the two-dimensional uncertainty prompt word module 400: The reference points, uncertainty, and BEV features transformed by the front view to bird's-eye view module 200 obtained by the front view uncertainty perception decoder 300 are used as the input of the two-dimensional uncertainty prompt word module 400. The reference points form an initial map vector through inverse perspective mapping (IPM). The initial map vector and uncertainty form a weight vector of the prompt word through a multilayer perceptron (MLP) and matrix multiplication. The weight vector of the prompt word is used by the front view to bird's-eye view attention module 402 and the front view transformation vector module 403 to obtain the updated BEV features and map vector respectively. In the figure, represents a quantity related to the uncertainty in the perspective view (PV). It may be an index used to quantify the uncertainty of certain elements in the perspective view (such as object positions, boundaries, etc.), and plays an important role in processes such as constructing perspective view prompt words, helping the system consider and handle information uncertainty. represents the weight of the perspective view (PV). This weight is related to the importance of various features and information in the perspective view. For example, in the process of constructing perspective view prompt words or subsequent feature fusion, is used to adjust the influence of perspective view information, determine the proportion of perspective view information in the overall processing flow, and thus affect the finally updated map vector and bird's-eye view features. is the average coordinate in the perspective view (PV). It can represent the average position coordinates of a certain area in the perspective view (such as a road area, an object area, etc.). When associating and fusing perspective view information with the map vector, etc., provides a position reference, which helps to accurately perform perspective conversion and feature alignment, ensuring that the information between the perspective view and the bird's-eye view can be correctly matched and interacted.
[0067] In an embodiment of the present invention, please refer to Figure 6 , The working principle of the bird's-eye view uncertainty perception decoder 500: The output of the two-dimensional uncertainty prompt word module 400 is input into the bird's-eye view uncertainty perception decoder 500, and then after passing through the self-attention module and normalization, it is input into the uncertainty perception attention module UA-attention to obtain the updated map vector. The updated map vector and the sampled image features are multiplied by a matrix to obtain the final map vector, and a high-precision map is reconstructed. In the figure, represents the reference vector, represents the mean value, represents the variance.
[0068] Please refer to Figure 8 , Figure 8 which shows a schematic diagram of the mechanism of the electronic device 20 in which embodiments of the present invention can be implemented. The electronic device is intended to represent various forms of control devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0069] The electronic device 20 includes at least one processor 21, and a memory communicatively connected to the at least one processor 21, such as a read-only memory (ROM) 22, a random access memory (RAM) 23, etc. The memory stores a computer program executable by the at least one processor. The processor 21 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 22 or the computer program loaded from the storage unit 28 into the random access memory (RAM) 13. In the RAM 23, various programs and data required for the operation of the electronic device 20 can also be stored. The processor 21, the ROM 22, and the RAM 23 are connected to each other via a bus 24. An input / output (I / O) interface 25 is also connected to the bus 24.
[0070] A plurality of components in the electronic device 20 are connected to the I / O interface 25, including: an input unit 26, such as a keyboard, a mouse, etc.; an output unit 27, such as various types of displays, speakers, etc.; a storage unit 28, such as a magnetic disk, an optical disc, etc.; and a communication unit 29, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 29 allows the electronic device 20 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0071] The processor 21 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 21 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 21 executes the various methods and processes described above.
[0072] In some embodiments, the method of the above embodiments may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 28. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 20 via the ROM 22 and / or the communication unit 29. When the computer program is loaded into the RAM 23 and executed by the processor 21, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the processor 21 may be configured to perform the methods of the above embodiments in any other suitable manner (e.g., by means of firmware).
[0073] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0074] The computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0075] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0076] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0077] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0078] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0079] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A high-precision map reconstruction method, characterized in that: The steps include: Extract the front view features of the surround image through the backbone network; Convert the front view feature into a bird's eye view feature; The object level information and point level information of the two-dimensional front view are obtained by decoding the front view features and the initialized map vector, and they are used as the front view prompt words. The decoding of the front view features and the initialized map vector specifically includes: Initialize the map vector and decode it together with the front view features; Through decoding calculation, the information of the initial reference point and uncertainty of the front view image is obtained, and this information constitutes a part of the front view prompt word; At the same time, the front view features are converted into bird's-eye view features; The front view prompt words, the initial map vector and the bird's-eye view features are used as two-dimensional uncertainty prompt words to generate updated map vectors and bird's-eye view features to further enrich the prompt word information; Combine the front view prompt words, the initialized map vector and the bird's-eye view features, and update the map vector and the bird's-eye view features; Get updated map vectors and bird's-eye view features to generate high-precision map reconstruction results.
2. A high-precision map reconstruction method according to claim 1, characterized in that: The surround view image refers to an image sequence of the surrounding environment collected by a surround view camera installed on a vehicle or other mobile platform, and obtained by preprocessing the image sequence.
3. A high-precision map reconstruction method according to claim 2, characterized in that: The image sequence preprocessing includes image denoising, graying, cropping, and scaling operations to improve image quality and make it suitable for input into the backbone network.
4. A high-precision map reconstruction method according to claim 3, characterized in that: The extracting of the front view features of the surround view image through the backbone network refers to inputting the preprocessed surround view image into the backbone network, and the backbone network extracts features of the image layer by layer according to its pre-trained model parameters and convolutional neural network structure, and finally outputs the front view features.
5. A high-precision map reconstruction system, used to implement the method according to any one of claims 1 to 4, characterized in that: The system includes: A backbone network, wherein the backbone network is used to extract front view features of the surround view image; A front view to bird's eye view module, which converts front view features into bird's eye view features; A front view uncertainty-aware decoder, wherein the front view uncertainty-aware decoder decodes the front view features and the initialized map vector to obtain two-dimensional front view object level information and point level information, and uses them as front view prompt words; A two-dimensional uncertainty prompt word module, wherein the two-dimensional uncertainty prompt word module combines the front view prompt word, the initialized map vector and the bird's-eye view feature, and updates the map vector and the bird's-eye view feature; A bird's-eye view uncertainty-aware decoder is provided, wherein the bird's-eye view uncertainty-aware decoder obtains updated map vectors and bird's-eye view features and generates a high-precision map reconstruction result.
6. An electronic device, characterized in that: The method comprises at least one processor, wherein the processor is communicatively connected to at least one memory, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method according to any one of claims 1 to 4 when executed.
Citation Information
Patent Citations
Map element detection method and device, electronic equipment and storage medium
CN116052097A
Spectrum map communication organization method, device and equipment based on semantic information
CN119006659A