Vehicle positioning method, neural network training method, and related device

By combining images and map information of the vehicle's surrounding environment, and using Transformer neural networks and multilayer perceptron modules to generate predictive information, the problem of insufficient vehicle positioning accuracy is solved, and high-precision positioning in complex environments is achieved.

WO2025180245A9PCT designated stage Publication Date: 2025-10-30YINWANG INTELLIGENT TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/077536
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-29
Filing Date
2025-02-17
Publication Date
2025-10-30

Smart Images

  • Figure CN2025077536_30102025_PF_FP_ABST
    Figure CN2025077536_30102025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a vehicle positioning method, a neural network training method, and a related device, in which the artificial intelligence technology can be applied to the field of intelligent driving. The method comprises: performing feature extraction on an image of a surrounding environment around a vehicle to obtain first feature information; performing feature extraction on map information in a preset range of a first location to obtain second feature information, wherein the first location is a location acquired by a positioning system deployed in the vehicle, and the map information comes from a standard-definition map; and on the basis of the first feature information and the second feature information, generating first prediction information by means of a first neural network, wherein the first prediction information is used for determining the actual location of the vehicle. The accuracy of the location (e.g. the location of a vehicle obtained by a GNSS) obtained by only a positioning system of the vehicle is low; therefore, in the present application, the location of the vehicle is finally determined on the basis of first location in combination with more visual information, so that the more accurate location of the vehicle can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

A vehicle localization method, a neural network training method, and related equipment.

[0001] This application claims priority to Chinese Patent Application No. 202410231541.9, filed on February 29, 2024, entitled "A method for locating a vehicle, a method for training a neural network and related equipment", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of intelligent driving, and in particular to a vehicle localization method, a neural network training method, and related equipment. Background Technology

[0003] In the field of intelligent driving, it is often necessary to locate vehicles. Specifically, vehicles can obtain their location information by using the Global Navigation Satellite System (GNSS). However, the accuracy of the vehicle location information obtained by GNSS is limited, meaning that there is often a large deviation between the vehicle location information obtained by GNSS and the actual location of the vehicle. Therefore, a more accurate positioning solution is urgently needed. Summary of the Invention

[0004] This application provides a vehicle positioning method, a neural network training method, and related equipment. Based on a first position, it integrates more visual information to finally determine the vehicle's position, which is beneficial for obtaining a more accurate vehicle position.

[0005] This application provides the following technical solution:

[0006] Firstly, this application provides a vehicle positioning method that can apply artificial intelligence technology to the field of intelligent driving. In this method, the execution device can acquire at least one image of the vehicle's surrounding environment, and after obtaining the vehicle's first position using the vehicle's positioning system, it can acquire map information within a preset range of the first position. The vehicle's positioning system may include at least one of the following: GNSS, a positioning system that uses a base station for positioning, or other types of positioning systems deployed in the vehicle.

[0007] The execution device extracts features from images of the vehicle's surrounding environment to obtain first feature information, and extracts features from map information within a preset range of the first location to obtain second feature information. The map information within the preset range of the first location is derived from a high-precision map. Based on the first and second feature information, the device generates first prediction information through a first neural network. The first prediction information is used to determine the vehicle's actual location.

[0008] For example, the precision of a standard-precision map can be lower than that of a high-precision map. For instance, the precision of a standard-precision map can be at the meter level, while the precision of a high-precision map can be at the centimeter level. For example, various navigation maps deployed on mobile phones are generally standard-precision maps, while the users of high-precision maps are generally computers. For instance, a standard-precision map only includes simple road lines, while a high-precision map carries detailed lane lines, road components, or other road information. A standard-precision map also includes information about buildings in the environment.

[0009] For example, in one scenario, the first prediction information indicates the predicted offset between the vehicle's actual position and a first position. After receiving the first prediction information generated by the first neural network, the vehicle can determine its actual position based on the first prediction information and the first position; this actual position is the actual position predicted by the neural network. In another scenario, the first prediction information indicates the vehicle's predicted actual position. For example, the first prediction information may include coordinate information corresponding to the vehicle's actual position. Optionally, the first prediction information may also include the vehicle's actual orientation.

[0010] Because the accuracy of relying solely on the vehicle's positioning system (such as GNSS-based vehicle location) is low in scenarios such as urban canyons, tunnels, or under overpasses with high-rise buildings, this implementation obtains a first location for the vehicle based on its positioning system. Then, it acquires map information within a preset range of the first location and extracts features from this map information to obtain second feature information. Furthermore, it extracts features from images of the vehicle's surrounding environment to obtain first feature information. Finally, it combines the first and second feature information to generate predictive information through a neural network. This predictive information is used to determine the vehicle's actual location. In other words, it integrates more visual information based on the first location to ultimately determine the vehicle's position, resulting in a more accurate location. Moreover, when this solution is applied to scenarios such as urban canyons or under overpasses with high-rise buildings, there will be buildings in the surrounding environment. The images of the surrounding environment will contain images of large buildings, and the high-resolution map will also contain building information. Therefore, obtaining map information within the preset range of the first location from a high-resolution map makes it easier to determine the vehicle's actual location.

[0011] In one possible implementation, the first neural network includes a Transformer neural network module and a multi-layer perceptron (MLP). The execution device generates first prediction information based on first feature information and second feature information through the first neural network. Specifically, this may include: the execution device fusing the first feature information and the second feature information to obtain fused feature information, and then inputting the fused feature information into the first neural network to obtain the first prediction information generated by the first neural network. The first prediction information indicates the offset between the vehicle's actual position and the first position.

[0012] For example, the first predicted information can use a first parameter and a second parameter to express the offset between the vehicle's actual position and the first position. In one case, the first parameter refers to the offset of the vehicle's actual position from the first position in the horizontal direction on the map. "Horizontal" can also be understood as the east-west direction, meaning the first parameter can represent the offset distance between the vehicle's actual position and the first position in the east-west direction. The first parameter also refers to the offset of the vehicle's actual position from the first position in the vertical direction on the map. "Vertical" can also be understood as the north-south direction, meaning the first parameter can represent the offset distance between the vehicle's actual position and the first position in the north-south direction. In another case, the first parameter refers to the offset of the vehicle's actual position from the first position in the direction perpendicular to the vehicle's front, and the second parameter refers to the offset of the vehicle's actual position from the first position in the direction the vehicle's front is pointing. In this case, the vehicle can decompose the first parameter into east-west and north-south directions, and the second parameter into east-west and north-south directions, thereby obtaining the offset distance between the vehicle's actual position and the first position in the east-west direction, and the offset distance between the vehicle's actual position and the first position in the north-south direction.

[0013] Optionally, the first prediction information may also include a third parameter, which refers to the offset between the vehicle's actual orientation and the first orientation. The first orientation can also be understood as the default orientation, which is often regarded as due north or due east. The "offset between the vehicle's actual orientation and the first orientation" can also be understood as the offset angle between the vehicle's actual orientation and due north, or the "offset between the vehicle's actual orientation and the first orientation" can also be understood as the offset angle between the vehicle's actual orientation and due east.

[0014] In this implementation, a specific network structure of the first neural network is provided (i.e., the first neural network includes a Transformer neural network module and a multilayer perceptron MLP), which reduces the implementation difficulty of this scheme. In addition, the Transformer neural network module and the multilayer perceptron MLP can generate the first prediction information relatively quickly, which helps to shorten the time delay of obtaining the actual position of the vehicle, that is, improve the efficiency of obtaining the actual position of the vehicle.

[0015] In one possible implementation, at least one image of the vehicle's surrounding environment may include at least two images of the vehicle's surrounding environment, which include images captured by the vehicle in at least two directions: left front, front, right front, left rear, rear, or right rear. In this implementation, acquiring images of the vehicle in multiple directions helps to more accurately reflect the vehicle's surrounding environment. Therefore, using images acquired by the vehicle in multiple directions to determine the vehicle's position information helps to make the final predicted actual position of the vehicle more accurate.

[0016] In one possible implementation, the first feature information includes feature information of the vehicle's surrounding environment from a top-down perspective. The execution device extracts features from the image of the vehicle's surrounding environment, which may include: the execution device extracting features from the image of the vehicle's surrounding environment using a second neural network. During the training of the second neural network using a first loss function (i.e., the training phase of the second neural network), the feature information generated by the second neural network is also input into a third neural network. The third neural network is used to generate second prediction information. The first loss function indicates the similarity between the second prediction information and the second expected information. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down perspective, and also indicates the predicted category of objects in the predicted image from the top-down perspective. The second expected information includes a correct image of the vehicle's surrounding environment from a top-down perspective, and also indicates the correct category of objects in the correct image from the top-down perspective. That is, both the second prediction information and the second expected information can be represented as images carrying labeled information, referring to images of the vehicle's surrounding environment from a top-down perspective. The difference is that the second prediction information is generated by the third neural network, and the second expected information is correct.

[0017] In this implementation, during the training phase of the second neural network, the feature information generated by the second neural network is input into the third neural network. The third neural network generates second prediction information, which indicates the predicted image of the vehicle's surrounding environment from a top-down perspective and the predicted categories of objects in the image. Then, the second neural network is trained using a first loss function, which indicates the similarity between the second prediction information and the second expected information. The second expected information includes the correct image of the vehicle's surrounding environment and the correct categories of objects in the image. That is, based on the feature information of the vehicle's surrounding environment generated by the second neural network from a top-down perspective, semantic segmentation of the vehicle's surrounding environment is performed. The correct semantic segmentation result is used as supervision to train the second neural network. Since the better the second feature information collected by the second neural network, the simpler the semantic segmentation of the third neural network is, and the easier it is to generate better quality semantic segmentation results, the higher the similarity between the second prediction information and the second expected information. Using this training method is beneficial to improving the second neural network's ability to better understand the vehicle's surrounding environment, which also helps to ensure that the first feature information extracted by the trained second neural network contains richer information.

[0018] In one possible implementation, the execution device may further rasterize the map information within a preset range of the first location in a first format to obtain map information within the preset range of the first location in a second format. The first format is either text or binary, and the second format is an image format. Then, feature extraction of the map information within the preset range of the first location by the execution device may include: feature extraction of the map information within the preset range of the first location in the second format.

[0019] The phrase "rasterizing the map information within a preset range of the first location in the first format" refers to performing a rendering operation based on the map information within the preset range of the first location in the first format to obtain a two-dimensional image-format map. During rasterization, a portion of the information included in the map information within the preset range of the first location in the first format is extracted, so that the map information within the preset range of the first location in the second format retains only a portion of the information from the map information within the preset range of the first location in the first format. Optionally, during rasterization, information of three types—area, lane, and node—can be extracted. That is, the rasterized map information can retain the shapes of objects within the preset range of the first location and the relative positional relationships between different objects.

[0020] In this implementation, when the map information acquired by the vehicle is in text or binary format, the first format map information can be rasterized to obtain the second format map information, which is an image format. Then, feature extraction is performed on the second format map information. Since the image format map information can more intuitively display the shape of objects within a preset range of the first location and the relative positions between different objects, feature extraction on the second format map information can more easily obtain rich image information. Moreover, after the rasterization process, only some important information in the map information within the preset range of the first location is retained, which is beneficial for extracting more critical information from the map information within the preset range of the first location. It is also beneficial for the second feature information to contain more important information, thereby enabling the first neural network to generate more accurate prediction information with the help of the first and second feature information.

[0021] In one possible implementation, the execution device extracts features from the map information within a preset range of the first location. This can include: the execution device directly inputting the acquired map information within the preset range of the first location into a fourth neural network, and extracting second feature information through the fourth neural network. Optionally, if the map information within the preset range of the first location acquired in step 302 is in a first format, the fourth neural network can also be a graph neural network (GNN), an attention-based neural network, or other types of neural networks.

[0022] In this implementation, since text-based map information includes the coordinates of multiple points and the category corresponding to each point, the text-based map information can be converted into a graph structure. Each vertex in the graph structure can be used to store the coordinates of a point and the category information corresponding to that point. In other words, the graph structure can effectively store text-based map information. Furthermore, the graph structure neural network can effectively extract features from the graph structure information. That is, when the graph neural network is used to extract features from the graph structure information, it can extract rich information from the graph structure information, thereby obtaining high-quality second feature information.

[0023] Secondly, this application provides a method for training a neural network, which can apply artificial intelligence technology to the field of intelligent driving. In this method, the training device inputs an image of the vehicle's surrounding environment into a second neural network, and extracts features from the image of the vehicle's surrounding environment through the second neural network to obtain first feature information; the first feature information is input into a third neural network to obtain second prediction information generated by the third neural network. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down view, and also indicates the predicted category of objects in the predicted image from the top-down view; the second neural network is trained using a first loss function to obtain a trained second neural network. The first loss function indicates the similarity between the second prediction information and the second expected information. The second expected information includes a correct image of the vehicle's surrounding environment from a top-down view, and also indicates the correct category of objects in the correct image from the top-down view.

[0024] In one possible implementation, the second expected information is obtained based on a standard-precision map and a high-precision map corresponding to the vehicle's surrounding environment. For example, the second expected information corresponding to the training samples includes a correct image of the vehicle's surrounding environment from a top-down view, containing roads and buildings; wherein, the road portion of the correct image of the vehicle's surrounding environment from a top-down view can be obtained based on a high-precision map, and the building portion can be obtained based on a standard-precision map; the correct categories of objects on the roads in the vehicle's surrounding environment can be obtained based on the high-precision map corresponding to the vehicle's surrounding environment, and the correct categories of buildings in the vehicle's surrounding environment can be obtained based on the standard-precision map corresponding to the vehicle's surrounding environment.

[0025] In this implementation, since the high-precision map carries detailed road information and the standard-precision map carries building information in the environment, and the second expected information includes a correct image of the vehicle's surrounding environment from a top-down perspective, and also indicates the correct category of objects in the correct image from a top-down perspective, using both the high-precision map and the standard-precision map to generate the second expected information is beneficial because the second expected information contains both high-quality road information and building information in the environment. In other words, the second expected information can more comprehensively and accurately reflect the vehicle's surrounding environment. Using the second expected information as a guide to train the second neural network is beneficial because the trained second neural network has a better understanding of the vehicle's surrounding environment. This is also beneficial because the first feature information extracted by the trained second neural network contains richer information, which in turn helps the first neural network generate more accurate prediction information with the help of the first and second feature information.

[0026] In one possible implementation, the training device can further extract features from map information within a preset range of the first location to obtain second feature information. The first location represents the vehicle's position obtained by a vehicle-based positioning system. Based on the first and second feature information, a first prediction information is generated through a first neural network. This first prediction information is used to determine the predicted actual position of the vehicle. The training device then trains a second neural network using a first loss function to obtain a trained second neural network. This can include: the training device using both the first and second loss functions to train the first and second neural networks, resulting in a trained first neural network and a trained second neural network. The second loss function indicates the similarity between the first prediction information and the first expected information, which is used to determine the correct actual position of the vehicle.

[0027] In the second aspect of this application, the training device may also perform the steps performed by the execution device in the various possible implementations of the first aspect. For the meanings of the terms in the second aspect of this application and the various possible implementations of the second aspect, as well as the beneficial effects brought about by each possible implementation, please refer to the descriptions in the various possible implementations of the first aspect, which will not be repeated here.

[0028] Thirdly, this application provides a vehicle positioning device that can apply artificial intelligence technology to the field of intelligent driving. The vehicle positioning device includes: an acquisition module for acquiring images of the vehicle's surrounding environment and map information within a preset range of a first location, wherein the first location is the vehicle's position obtained based on a vehicle positioning system; a feature extraction module for extracting features from the images of the vehicle's surrounding environment to obtain first feature information; the feature extraction module is further used for extracting features from the map information within the preset range of the first location to obtain second feature information; and a generation module for generating first prediction information based on the first and second feature information using a first neural network, wherein the first prediction information is used to determine the vehicle's actual position.

[0029] In one possible implementation, the first neural network includes a Transformer neural network module and a multilayer perceptron (MLP), and a generation module specifically used for: fusing first feature information and second feature information to obtain fused feature information; inputting the fused feature information into the first neural network to obtain first prediction information generated by the first neural network, wherein the first prediction information indicates the offset between the actual position of the vehicle and the first position.

[0030] In one possible implementation, the image of the vehicle's surrounding environment includes at least two images of the vehicle's surrounding environment, which include images acquired by the vehicle in at least two directions: left front, front, right front, left rear, rear, or right rear.

[0031] In one possible implementation, the first feature information includes feature information of the vehicle's surrounding environment from a top-down perspective. A feature extraction module is specifically used to extract features from the image of the vehicle's surrounding environment using a second neural network. During the training of the second neural network using a first loss function, the feature information generated by the second neural network is also input into a third neural network. The third neural network is used to generate second prediction information. The first loss function indicates the similarity between the second prediction information and the second expected information. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down perspective, and also indicates the predicted category of objects in the predicted image from the top-down perspective. The second expected information includes a correct image of the vehicle's surrounding environment from a top-down perspective, and also indicates the correct category of objects in the correct image from the top-down perspective.

[0032] In one possible implementation, the vehicle positioning device further includes: a rasterization module, used to rasterize map information within a preset range of a first location in a first format to obtain map information within a preset range of a first location in a second format, wherein the first format is a text format or a binary format, and the second format is an image format; and a feature extraction module, specifically used to extract features from the map information within a preset range of the first location in the second format.

[0033] The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the third aspect can all be found in the first aspect, and will not be repeated here.

[0034] Fourthly, this application provides a neural network training device that can apply artificial intelligence technology to the field of intelligent driving. The neural network training device includes: an input module for inputting an image of the vehicle's surrounding environment into a second neural network, and extracting features from the image of the vehicle's surrounding environment through the second neural network to obtain first feature information; a generation module for inputting the first feature information into a third neural network to obtain second prediction information generated by the third neural network, the second prediction information including a predicted image of the vehicle's surrounding environment from a top-down view, and the second prediction information also indicating the predicted category of objects in the predicted image from the top-down view; and a training module for training the second neural network using a first loss function to obtain a trained second neural network, the first loss function indicating the similarity between the second prediction information and second expected information, the second expected information including a correct image of the vehicle's surrounding environment from a top-down view, and the second expected information also indicating the correct category of objects in the correct image from the top-down view.

[0035] In one possible implementation, the second desired information is obtained based on a standard-precision map and a high-precision map corresponding to the vehicle's surrounding environment.

[0036] In one possible implementation, the training device for the neural network further includes: a feature extraction module, used to extract features from map information within a preset range of the first location to obtain second feature information, wherein the first location represents the location of the vehicle obtained by a vehicle-based positioning system; a generation module, further used to generate first prediction information through a first neural network based on the first feature information and the second feature information, wherein the first prediction information is used to determine the predicted actual location of the vehicle; and a training module, specifically used to train the first neural network and the second neural network using a first loss function and a second loss function to obtain the trained first neural network and the trained second neural network, wherein the second loss function indicates the similarity between the first prediction information and the first expected information, wherein the first expected information is used to determine the correct actual location of the vehicle.

[0037] In the fourth aspect of this application, the training device for the neural network is also used to perform the steps executed by the training device in the second aspect and various possible implementations of the second aspect. The specific implementation methods, the meanings of the terms, and the beneficial effects of the steps in the fourth aspect can be found in the second aspect, and will not be repeated here.

[0038] Fifthly, embodiments of this application provide an apparatus including a processor and a memory, the processor being coupled to the memory, the memory being used to store a program; the processor being used to execute the program in the memory, causing the apparatus to perform the methods described in the first or second aspect above.

[0039] In a sixth aspect, embodiments of this application provide a vehicle including a processor and a memory, the processor being coupled to the memory, the memory being used to store a program; and the processor being used to execute the program in the memory, causing the vehicle to perform the method described in the first aspect above.

[0040] In a seventh aspect, embodiments of this application provide a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to perform the methods described in the first or second aspect above.

[0041] Eighthly, embodiments of this application provide a computer program product, which includes a program that, when run on a computer, causes the computer to perform the methods described in the first or second aspect above.

[0042] Ninthly, this application provides a chip system including a processor for supporting the implementation of the functions involved in the foregoing aspects, such as transmitting or processing data and / or information involved in the foregoing methods. In one possible design, the chip system further includes a memory for storing program instructions and data necessary for a terminal device or communication device. This chip system may be composed of chips or may include chips and other discrete devices.

[0043] The second to eighth aspects of this application correspond to multiple possible ways of the first aspect and have corresponding beneficial effects. Attached Figure Description

[0044] Figure 1 is a schematic diagram of a structural framework for an artificial intelligence main body provided in an embodiment of this application;

[0045] Figure 2 is a schematic diagram of an architecture of a vehicle positioning system provided in an embodiment of this application;

[0046] Figure 3 is a flowchart illustrating a vehicle positioning method provided in an embodiment of this application;

[0047] Figure 4 is a schematic diagram of an image of the vehicle's surrounding environment obtained according to an embodiment of this application;

[0048] Figure 5 is a schematic diagram of extracting features from an image of the vehicle's surrounding environment to obtain the first feature information;

[0049] Figure 6 is a schematic diagram of extracting features from map information within a preset range of a first location to obtain second feature information, according to an embodiment of this application.

[0050] Figure 7 is a schematic diagram of the image of the vehicle's surrounding environment, first feature information, second feature information, and the actual position of the vehicle provided in an embodiment of this application.

[0051] Figure 8 is a schematic diagram of generating first prediction information through a first neural network according to an embodiment of this application;

[0052] Figure 9 is a schematic diagram of another vehicle positioning method provided in an embodiment of this application;

[0053] Figure 10 is a flowchart illustrating a neural network training method provided in an embodiment of this application;

[0054] Figure 11 is a schematic diagram of another neural network training method provided in the embodiments of this application;

[0055] Figure 12 is a schematic diagram of a vehicle positioning device provided in an embodiment of this application;

[0056] Figure 13 is a schematic diagram of a neural network training device provided in an embodiment of this application;

[0057] Figure 14 is a schematic diagram of a device provided in an embodiment of this application;

[0058] Figure 15 is a schematic diagram of another structure of the device provided in an embodiment of this application;

[0059] Figure 16 is a structural schematic diagram of a vehicle provided in an embodiment of this application. Detailed Implementation

[0060] The embodiments of this application will now be described with reference to the accompanying drawings. Obviously, the described embodiments are merely some, and not all, of the embodiments of this application. Those skilled in the art will recognize that, with the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0061] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the description of embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of units is not necessarily limited to those units, but may include other units not explicitly listed or inherent to those processes, methods, products, or apparatuses.

[0062] In the embodiments of this application, "instruction" can include direct and indirect instructions, as well as explicit and implicit instructions. The information indicated by a certain piece of information (hereinafter referred to as instruction information) is called the information to be instructed. In specific implementation, there are many ways to indicate the information to be instructed, such as, but not limited to, directly indicating the information to be instructed, such as the information to be instructed itself or its index. It can also indirectly indicate the information to be instructed by indicating other information, where there is an association between the other information and the information to be instructed; or it can indicate only a part of the information to be instructed, while the other parts are known or pre-agreed upon. For example, the instruction can be implemented by using a pre-agreed (e.g., protocol predefined) arrangement of various information, thereby reducing the instruction overhead to a certain extent. This application does not limit the specific method of instruction. It is understood that for the sender of the instruction information, the instruction information can be used to indicate the information to be instructed; for the receiver of the instruction information, the instruction information can be used to determine the information to be instructed.

[0063] First, the overall workflow of an artificial intelligence system is described, as shown in Figure 1. Figure 1 is a structural diagram of the main framework of artificial intelligence. The framework is then elaborated on from two dimensions: the "Intelligent Information Chain" (horizontal axis) and the "IT Value Chain" (vertical axis). The "Intelligent Information Chain" reflects a series of processes from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. In this process, data undergoes a condensation process of "data—information—knowledge—wisdom." The "IT Value Chain" reflects the value that artificial intelligence brings to the information technology industry, from the underlying infrastructure of human intelligence and information (provided and processed by technology) to the industrial ecosystem of the system.

[0064] (1) Infrastructure

[0065] The infrastructure provides computing power to support artificial intelligence systems, enabling communication with the external world and providing support through a basic platform. Communication with the outside world is achieved through sensors; computing power is provided by intelligent chips, which can specifically employ hardware acceleration chips such as central processing units (CPUs), embedded neural network processing units (NPUs), graphics processing units (GPUs), application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The basic platform includes distributed computing frameworks and related platform guarantees and support, which may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside world to acquire data, and this data is provided to intelligent chips in the distributed computing system provided by the basic platform for computation.

[0066] (2) Data

[0067] The data at the next layer of infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voice, text, and IoT data from traditional devices, including business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.

[0068] (3) Data processing

[0069] Data processing typically includes methods such as data training, machine learning, deep learning, search, reasoning, and decision-making.

[0070] Among them, machine learning and deep learning can perform intelligent information modeling, extraction, preprocessing, and training on data, including symbolization and formalization.

[0071] Reasoning refers to the process in which, in a computer or intelligent system, the machine thinks and solves problems by simulating human intelligent reasoning, based on reasoning control strategies and using formalized information. Typical functions include search and matching.

[0072] Decision-making refers to the process of making decisions based on intelligent information after reasoning, and it typically provides functions such as classification, sorting, and prediction.

[0073] (4) General ability

[0074] After the data processing mentioned above, the results of the data processing can be used to form some general capabilities, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0075] (5) Smart Products and Industry Applications

[0076] Intelligent products and industry applications refer to products and applications of artificial intelligence systems in various fields. They encapsulate overall artificial intelligence solutions, productize intelligent information decision-making, and realize practical applications. Their application areas mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, intelligent driving, and smart cities.

[0077] The method provided in this application can be applied to various scenarios requiring vehicle positioning. For example, in the field of intelligent driving, where vehicles can provide navigation functions, navigation routes need to be planned based on the vehicle's actual location. The vehicles can be cars, trucks, motorcycles, buses, boats, lawnmowers, recreational vehicles, amusement park vehicles, construction equipment, trams, golf carts, trains, airplanes, and helicopters, etc., and this application does not impose any particular limitation.

[0078] Currently, the location of a vehicle can be obtained using the Global Navigation Satellite System (GNSS) deployed inside the vehicle. However, the location information obtained by GNSS often deviates significantly from the actual location of the vehicle. Therefore, a more accurate positioning solution is urgently needed.

[0079] To address the aforementioned issues, this application provides a vehicle positioning method. The method discloses that after obtaining a first location of the vehicle based on a vehicle positioning system, map information within a preset range of the first location can be acquired. Feature extraction is then performed on the map information within this preset range to obtain second feature information. Furthermore, feature extraction is performed on images of the vehicle's surrounding environment to obtain first feature information. Finally, the first and second feature information are combined using a neural network to generate prediction information. This prediction information is used to determine the vehicle's actual location. In other words, based on the first location obtained from the positioning system, more visual information is integrated to ultimately determine the vehicle's location, resulting in a more accurate vehicle position.

[0080] Before providing a detailed description of the method provided in this application, the architecture of the vehicle positioning system provided in this application will be described first. Please refer to Figure 2, which is a schematic diagram of the architecture of a vehicle positioning system provided in an embodiment of this application. As shown in Figure 2, the vehicle positioning system 200 includes a training device 210, a database 220, an execution device 230, and a data storage system 240. The execution device 230 includes a computing module 231.

[0081] The database 220 stores a training dataset. During the training phase of the first neural network 201, the training device 210 generates the first neural network 201 and iteratively trains it using the training dataset to obtain a first neural network 201 that has undergone training operations. The "first neural network 201 that has undergone training operations" can also be called the "trained first neural network 201". The first neural network 201 can be specifically represented as a neural network or as a non-neural network model. In this embodiment, the first neural network 201 is only described as a neural network.

[0082] The first neural network 201, trained by the training device 210, can be deployed to the computing module 231 of the execution device 230. The execution device 230 can access data, code, etc., in the data storage system 240, and can also store data, instructions, etc., in the data storage system 240. The data storage system 240 can be located within the execution device 230, or it can be an external memory relative to the execution device 230.

[0083] During the application phase of the first neural network 201, the execution device 230 can determine whether the vehicle veers off course when passing through a fork in the road based on images or point cloud data of the vehicle's surrounding environment through the trained first neural network 201.

[0084] In some embodiments of this application, referring to FIG2, the execution device 230 and the client device can be integrated into the same device, so the user can directly interact with the execution device 230. For example, when the client device is a vehicle, the execution device 230 can be a module in the vehicle's host CPU that uses a first neural network to process data. The execution device 230 can also be a graphics processing unit (GPU) or a neural network processor (NPU) in the vehicle. The GPU or NPU is mounted on the host processor as a coprocessor, and the host processor allocates tasks.

[0085] It is worth noting that Figure 2 is merely a schematic diagram of one architecture of the vehicle positioning system provided in an embodiment of the present invention, and the positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in some other embodiments of this application, the execution device 230 and the client device can be separate independent devices. The execution device 230 is configured with an input / output (I / O) interface to interact with the client device. The client device sends images of the vehicle's surrounding environment to the execution device 230 through the I / O interface. After determining the actual position of the vehicle through the first neural network 201 in the calculation module 231, the execution device 230 can return the actual position of the vehicle to the client device through the I / O interface.

[0086] Based on the above description, the specific implementation process of the training and application phases in the method provided in this application will now be described.

[0087] I. Application Phase

[0088] In this embodiment, the application stage refers to the process by which the execution device uses the trained first neural network to determine the actual location of the vehicle. The execution device can specifically be a vehicle or a processor within a vehicle. Subsequent embodiments will only use a vehicle as an example for illustration; when the execution device takes on other product forms, the same principle applies, and will not be elaborated upon in this embodiment. Specifically, please refer to Figure 3, which is a flowchart illustrating a vehicle positioning method provided in this embodiment. The vehicle positioning method provided in this embodiment may include:

[0089] 301. Obtain images of the environment surrounding the vehicle.

[0090] In this embodiment of the application, the vehicle's location information can be determined by using at least one image of the vehicle's surrounding environment. Optionally, the aforementioned image of the vehicle's surrounding environment includes at least two images of the vehicle's surrounding environment, which include images collected by the vehicle in at least two directions: left front, front, right front, left rear, rear, right rear, or other directions. The specific directions in which the vehicle's images are acquired can be flexibly determined based on the actual situation, and this application does not impose any limitations.

[0091] To understand this solution more intuitively, please refer to Figure 4. Figure 4 is a schematic diagram of an image of the vehicle's surrounding environment obtained according to an embodiment of this application. The image of the vehicle's surrounding environment obtained in Figure 4 includes, for example, images of the vehicle's left front, front, right front, left rear, rear, and right rear. That is, the image of the vehicle's surrounding environment obtained includes a surround view image of the vehicle. It should be understood that the example in Figure 4 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0092] In this embodiment of the application, acquiring images of the vehicle from multiple directions helps to more accurately reflect the vehicle's surrounding environment. Therefore, using images of the vehicle from multiple directions to determine the vehicle's position information helps to make the final predicted actual position of the vehicle more accurate.

[0093] 302. Obtain map information within a preset range of the first location. The first location is the vehicle's location obtained from the vehicle's positioning system, and the map information within the preset range of the first location comes from a high-precision map.

[0094] In this embodiment of the application, when it is necessary to obtain the vehicle's location information, the vehicle's first location can also be obtained based on the vehicle's positioning system, and then map information within a preset range of the first location (hereinafter referred to as the "first preset range" for easy distinction) can be obtained.

[0095] For example, the vehicle's positioning system may include at least one of the following: GNSS, a positioning system that uses a base station for positioning, or other types of positioning systems deployed in the vehicle, etc., which are not limited in this application. "The vehicle's first position" can also be understood as the initial position of the vehicle obtained based on the vehicle's positioning system.

[0096] For example, a navigation map may be deployed in the vehicle. After the vehicle determines the first location, it can obtain map information within a first preset range of the first location based on the navigation map. For example, the navigation map may be a standard-definition map (SD Map), a high-definition map (HD Map), or other types of maps.

[0097] For example, the precision of a standard-precision map can be lower than that of a high-precision map. For instance, the precision of a standard-precision map can be at the meter level, while the precision of a high-precision map can be at the centimeter level. For example, various navigation maps deployed on mobile phones are generally standard-precision maps, while the users of high-precision maps are generally computers. For instance, a standard-precision map only includes simple road lines, while a high-precision map carries detailed lane lines, road components, or other road information. A standard-precision map also includes information about buildings in the environment.

[0098] For example, "the first preset range of the first position" can be represented as a circular area centered on the first position, and the radius of the circular area can be a first preset length; for example, the first preset length can be 40 meters, 50 meters, 60 meters, 80 meters, 100 meters, or other values, etc., which are not limited in this application. As another example, "the first preset range of the first position" can be represented as a square area centered on the first position, and the side length of the square area can be a second preset length, which can be 80 meters, 100 meters, 120 meters, 160 meters, 200 meters, or other values, etc., which are not limited in this application.

[0099] 303. Extract features from the image of the vehicle's surrounding environment to obtain the first feature information.

[0100] In this embodiment of the application, when it is necessary to obtain the vehicle's location information, the vehicle can extract features from the image of the vehicle's surrounding environment obtained in step 301 through a second neural network to obtain the first feature information.

[0101] For example, the second neural network can be an encoder. Specifically, the second neural network can be a convolutional neural network (CNN), a residual neural network, or a neural network based on an attention mechanism. The specific neural network used can be flexibly determined based on the actual application scenario.

[0102] Optionally, the first feature information includes feature information of the vehicle's surrounding environment from a top-down perspective. "Feature information of the vehicle's surrounding environment from a top-down perspective" can also be referred to as "feature information of the vehicle's surrounding environment from a bird's-eye view perspective".

[0103] For example, in step 303, the vehicle can extract features from each image of the vehicle's surrounding environment through the second neural network to obtain the third feature information of each image. The third feature information refers to the feature information obtained without viewpoint transformation. The "feature information obtained without viewpoint transformation" can also be understood as the feature information of the vehicle's surrounding environment under perspective. The third feature information is converted into the feature information of each image under top-down view through the second neural network. The feature information of all images of the vehicle's surrounding environment under top-down view is fused by the second neural network to obtain the first feature information.

[0104] Alternatively, after obtaining the third feature information of each image, the vehicle can, based on the intrinsic parameters of the camera capturing each image of the vehicle's surrounding environment, convert the third feature information of each image into feature information in the camera coordinate system using a second neural network; then, based on the aforementioned extrinsic parameters of the camera, the second neural network can perform geometric projection of the feature information in the camera coordinate system to obtain the feature information of each image from a top-down perspective; finally, the second neural network can fuse the feature information of all images of the vehicle's surrounding environment from a top-down perspective to obtain the first feature information. For example, the extrinsic parameters of the camera may include the camera's installation position within the vehicle.

[0105] Alternatively, after obtaining the third feature information of each image, the vehicle can also use the intrinsic and extrinsic parameters of the camera that collects each image of the vehicle's surrounding environment to directly convert the third feature information of each image (that is, the feature information of each image under the perspective view) into the feature information of each image under the top view view through the second neural network, and then fuse the feature information of all images of the vehicle's surrounding environment under the top view view to obtain the first feature information.

[0106] To more intuitively understand this solution, please refer to Figure 5. Figure 5 is a schematic diagram of extracting features from images of the vehicle's surrounding environment to obtain the first feature information. The images of the vehicle's surrounding environment obtained in Figure 5 include, for example, six images from the vehicle's left front, front, right front, left rear, rear, and right rear. As shown in Figure 5, features can be extracted from each image of the vehicle's surrounding environment using a second neural network to obtain the third feature information for each image. Then, the second neural network converts the third feature information of each image into feature information from a top-down view. Finally, the feature information from the six images of the vehicle's surrounding environment from the top-down view is fused to obtain the first feature information. It should be understood that the first feature information shown in Figure 5 is an image obtained after visualizing the first feature information. It should also be understood that the example in Figure 5 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0107] Optionally, during the training of the second neural network using the first loss function, i.e., the training phase of the second neural network, the feature information generated by the second neural network is also input into the third neural network. The third neural network is used to generate second prediction information. The aforementioned first loss function indicates the similarity between the second prediction information and the second expected information. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down view, and the second expected information also indicates the predicted category of objects in the predicted image from the top-down view. The second expected information also includes a correct image of the vehicle's surrounding environment from a top-down view, and the second expected information also indicates the correct category of objects in the correct image from the top-down view. It should be noted that the specific implementation of the "training phase of the second neural network" will be described in detail in subsequent embodiments and will not be elaborated here.

[0108] Optionally, the vehicle can also acquire at least one point cloud data of the vehicle's surrounding environment. Optionally, the aforementioned point cloud data of the vehicle's surrounding environment includes at least two point cloud data of the vehicle's surrounding environment. The at least two point cloud data of the vehicle's surrounding environment include point cloud data collected by the vehicle in at least two directions: left front, front, right front, left rear, front rear, right rear, or other directions, etc., which are not limited in this embodiment. Then, in step 303, the vehicle can also use another neural network to extract features from at least one point cloud data of the vehicle's surrounding environment to obtain the feature information of the aforementioned point cloud data. After the vehicle extracts features from the images of the vehicle's surrounding environment through the second neural network, it can obtain the feature information of each image of the vehicle's surrounding environment from a top-down perspective. After fusing the feature information of all images of the vehicle's surrounding environment from a top-down perspective, it can obtain the fused feature information. After further fusing the feature information of the point cloud data and the fused feature information, the vehicle can obtain the first feature information.

[0109] 304. Extract features from the map information within the first preset range of the first location to obtain the second feature information.

[0110] In this embodiment of the application, when it is necessary to obtain the vehicle's location information, after the vehicle obtains the map information within the first preset range of the first location, it can also extract features from the map information within the first preset range of the first location through a fourth neural network to obtain second feature information. The second feature information includes the feature information of the map information within the first preset range of the first location.

[0111] For example, the fourth neural network can be an encoder. For instance, the fourth neural network can specifically be a convolutional neural network, a residual neural network, an attention-based neural network, a graph neural network, etc. The specific type can be determined based on the actual application scenario, and no limitation is made in this embodiment.

[0112] In one scenario, the map information within the first preset range of the first location obtained in step 302 is in a first format, which can be text, binary, or other formats. The text-formatted map information can describe the objects in the environment and their locations. For example, the text-formatted map information may include the coordinates of multiple points and the category corresponding to each point. For instance, the category corresponding to each point could be a building, traffic light pole, lane line, or other categories. Optionally, if the navigation map is a high-precision map, the text-formatted map information may also include the number of lanes. In another scenario, the map information within the first preset range of the first location obtained in step 302 can be in image format.

[0113] If the vehicle obtains map information within a first preset range of a first location in a first format, optionally, before executing step 304, the map information within the first preset range of the first location in the first format can be rasterized to obtain map information within the first preset range of the first location in a second format, where the second format is an image format. In one implementation, step 304 may include: the vehicle extracting features from the map information within the first preset range of the first location in the second format using a fourth neural network.

[0114] "Rasterizing the map information within the first preset range of the first location in the first format" refers to performing a rendering operation based on the map information within the first preset range of the first location in the first format to obtain map information in a two-dimensional image format.

[0115] During the rasterization process, a portion of the information included in the map information within the first preset range of the first location in the first format is extracted. Thus, the map information within the first preset range of the first location in the second format retains only a portion of the information from the first preset range of the first location in the first format. Optionally, during rasterization, information of three types—area, lane, and node—can be extracted. This means that the rasterized map information can retain the shapes of objects within the first preset range of the first location and the relative positional relationships between different objects.

[0116] To more intuitively understand this solution, please refer to Figure 6. Figure 6 is a schematic diagram of a method for extracting features from map information within a first preset range at a first location to obtain second feature information, as provided in this embodiment of the application. As shown in Figure 6, after obtaining map information within a first preset range at a first location in a first format, the vehicle first rasterizes the map information within the first preset range at the first location in the first format to obtain map information within the first preset range at the first location in a second format. Figure 6 also shows map information obtained after visualizing the map information within the first preset range at the first location in the first format. Comparing the map information in the first format (the image obtained after visualization) with the map information in the second format, it can be seen that only the shapes of objects within the first preset range at the first location and the relative positional relationships between different objects are retained during the rasterization process. Then, the vehicle extracts features from the map information within the first preset range at the first location in the second format using a fourth neural network to obtain second feature information. Figure 6 shows a schematic diagram obtained after visualizing the second feature information. It should be noted that the second feature information shown in Figure 6 is an image obtained after visualizing the second feature information. The example in Figure 6 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0117] In this embodiment, when the map information acquired by the vehicle is in text or binary format, the first format map information can be rasterized to obtain the second format map information, which is an image format. Then, feature extraction is performed on the second format map information. Since the image format map information can more intuitively display the shape of objects within the first preset range of the first location and the relative positions between different objects, feature extraction on the second format map information can more easily obtain rich image information. Moreover, after the rasterization process, only some important information in the map information within the first preset range of the first location is retained, which is beneficial for extracting more critical information in the map information within the first preset range of the first location. It is also beneficial for the second feature information to contain more important information, thereby enabling the first neural network to generate more accurate prediction information with the help of the first and second feature information.

[0118] In another implementation, step 304 may include: directly inputting the map information of the first location within a first preset range obtained in step 302 into the fourth neural network, and extracting the second feature information through the fourth neural network.

[0119] Optionally, if the map information within the first preset range of the first location obtained in step 302 is in the first format, the fourth neural network can also employ a graph neural network (GNN), an attention-based neural network, or other types of neural networks. Since text-based map information includes the coordinates of multiple points and the category corresponding to each point, it can be converted into a graph structure. Each vertex in the graph structure can store the coordinates of a point and the category information corresponding to that point. For example, the category information corresponding to the point can be the embedding obtained by vectorizing the category corresponding to the point. In other words, the graph structure effectively stores text-based map information, and the graph neural network can effectively extract features from the graph structure information. That is, when using a graph neural network to extract features from the graph structure information, rich information can be extracted, resulting in high-quality second feature information.

[0120] 305. Based on the first feature information and the second feature information, a first prediction information is generated through a first neural network, and the first prediction information is used to determine the actual position of the vehicle.

[0121] In this embodiment of the application, for example, the first neural network may include a Transformer neural network module and a multi-layer perceptron (MLP), or the first neural network may include a convolutional neural network and an MLP, or the first neural network may include a graph neural network and an MLP, or the first neural network may specifically adopt an MLP, or the first neural network may also be specifically represented as other network structures, which are not limited in this embodiment of the application.

[0122] In one scenario, the first prediction information indicates the predicted offset between the vehicle's actual position and the first position.

[0123] For example, the first predicted information can use a first parameter and a second parameter to express the offset between the vehicle's actual position and the first position. In one case, the first parameter refers to the offset of the vehicle's actual position from the first position in the horizontal direction on the map. "Horizontal" can also be understood as the east-west direction, meaning the first parameter can represent the offset distance between the vehicle's actual position and the first position in the east-west direction. The first parameter also refers to the offset of the vehicle's actual position from the first position in the vertical direction on the map. "Vertical" can also be understood as the north-south direction, meaning the first parameter can represent the offset distance between the vehicle's actual position and the first position in the north-south direction. In another case, the first parameter refers to the offset of the vehicle's actual position from the first position in the direction perpendicular to the vehicle's front, and the second parameter refers to the offset of the vehicle's actual position from the first position in the direction the vehicle's front is pointing. In this case, the vehicle can decompose the first parameter into east-west and north-south directions, and the second parameter into east-west and north-south directions, thereby obtaining the offset distance between the vehicle's actual position and the first position in the east-west direction, and the offset distance between the vehicle's actual position and the first position in the north-south direction.

[0124] Optionally, the first prediction information may also include a third parameter, which refers to the offset between the vehicle's actual orientation and the first orientation. The first orientation can also be understood as the default orientation, which is often regarded as due north or due east. The "offset between the vehicle's actual orientation and the first orientation" can also be understood as the offset angle between the vehicle's actual orientation and due north, or the "offset between the vehicle's actual orientation and the first orientation" can also be understood as the offset angle between the vehicle's actual orientation and due east.

[0125] To more intuitively understand this solution, please refer to Figure 7. Figure 7 is a schematic diagram of the image of the vehicle's surrounding environment, first feature information, second feature information, and the actual position of the vehicle provided in this application embodiment. Figure 7 includes four sub-schematic diagrams: (a), (b), (c), and (d). Sub-schematic diagram (a) of Figure 7 represents the image of the vehicle's surrounding environment obtained in step 301, which includes four images. Sub-schematic diagram (b) of Figure 7 represents the image obtained after visualizing the first feature information, that is, the image obtained after visualizing the feature information of the vehicle's surrounding environment from a top-down perspective. Sub-schematic diagram (c) of Figure 7 represents the image obtained after visualizing the second feature information, that is, the image obtained after visualizing the feature information of the map information within a first preset range of the first position. Sub-schematic diagram (c) of Figure 7 also shows a box and a black dot. The black dot represents the center point of the vehicle's surrounding environment, and the box represents the range obtained by mapping the vehicle's surrounding environment to the first preset range of the first position. In the schematic diagram of Figure 7(d), the large gray box represents the first preset range of the first position, the small gray box represents the range of the vehicle's surrounding environment, the gray dot in the lower left corner represents the vehicle's first position, the vehicle's initial direction is due north, and this gray dot is located in the lower left corner of the white dashed box; the white dot represents the vehicle's predicted actual position obtained through neural network prediction, and the white dashed box represents the vehicle's predicted orientation. It should be understood that the example in Figure 7 is only for the convenience of understanding this scheme and is not intended to limit this scheme.

[0126] Specifically, in one implementation, step 305 may include: the vehicle fusing the first feature information and the second feature information to obtain fused feature information, and then inputting the fused feature information into a first neural network to obtain first prediction information generated by the first neural network, the first prediction information indicating the offset between the vehicle's actual position and the first position; the vehicle determining its actual position based on the first prediction information and the first position, the actual position being the actual position predicted by the neural network.

[0127] To understand this solution more intuitively, please refer to Figure 8. Figure 8 is a schematic diagram of generating first prediction information through a first neural network according to an embodiment of this application. As shown in Figure 8, the first neural network includes a transformer module and a multilayer perceptron (MLP). After the vehicle obtains the first feature information and the second feature information, it first fuses the first feature information and the second feature information to obtain fused feature information. The fused feature information is then input into the first neural network to obtain the first prediction information output by the first neural network. The first prediction information indicates the offset between the actual position of the vehicle and the first position. It should be understood that the example in Figure 9 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0128] In this embodiment of the application, a specific network structure of the first neural network is provided (that is, the first neural network includes a Transformer neural network module and a multilayer perceptron MLP), which reduces the implementation difficulty of this solution; in addition, the first prediction information can be generated relatively quickly by using the Transformer neural network module and the multilayer perceptron MLP, which helps to shorten the time delay of obtaining the actual position of the vehicle, that is, to improve the efficiency of obtaining the actual position of the vehicle.

[0129] In another implementation, the vehicle can also directly input the first feature information and the second feature information into the first neural network to obtain the first prediction information generated by the first neural network. The first prediction information indicates the offset between the vehicle's actual position and the first position. The vehicle determines its actual position based on the first prediction information and the first position. This actual position is the actual position predicted by the neural network.

[0130] To more intuitively understand this solution, please refer to Figure 9. Figure 9 is a schematic diagram of a vehicle positioning method provided in an embodiment of this application. As shown in Figure 9, the vehicle can obtain its first position through a positioning system deployed on the vehicle; the vehicle can extract features from the image of the surrounding environment of the vehicle through a second neural network to obtain first feature information; a fourth neural network can extract features from the map information within a first preset range of the first position to obtain second feature information; based on the first feature information and the second feature information, a first prediction information is generated through a first neural network. In Figure 9, the first prediction information indicates the offset between the vehicle's first position and its actual position as an example; the vehicle can finally determine its actual position based on the first position and the first prediction information. It should be understood that the example in Figure 9 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0131] In another scenario, the first prediction information indicates the predicted actual position of the vehicle. For example, the first prediction information may include coordinate information corresponding to the vehicle's actual position. Optionally, the first prediction information may also include the vehicle's actual orientation. In this case, step 305 may include: the vehicle generating the first prediction information using a first neural network based on the first feature information, the second feature information, and the first position.

[0132] Specifically, in one implementation, the vehicle can fuse the first feature information and the second feature information to obtain fused feature information; the fused feature information and the first position are then input into a first neural network to obtain first prediction information generated by the first neural network. In another implementation, the vehicle can directly input the first feature information, the second feature information, and the first position into the first neural network to obtain first prediction information generated by the first neural network.

[0133] Because the accuracy of relying solely on vehicle positioning systems (such as GNSS-based vehicle positions) is low in scenarios such as urban canyons, tunnels, or under overpasses with high-rise buildings, this embodiment of the application obtains a first vehicle position based on the vehicle positioning system. Then, map information within a preset range of the first position is acquired, and features are extracted from this map information to obtain second feature information. Furthermore, features are extracted from images of the vehicle's surrounding environment to obtain first feature information. Finally, the first and second feature information are combined using a neural network to generate prediction information. This prediction information is used to determine the vehicle's actual position. In other words, more visual information is integrated on top of the first position to ultimately determine the vehicle's position, leading to a more accurate location. Moreover, when this solution is applied to scenarios such as urban canyons or under overpasses with high-rise buildings, buildings exist in the surrounding environment. Therefore, images of the surrounding environment will contain images of large buildings, and the high-resolution map will also contain building information. Thus, map information within the preset range of the first position derived from the high-resolution map makes it easier to determine the vehicle's actual position.

[0134] II. Training Phase

[0135] Specifically, please refer to Figure 10. Figure 10 is a flowchart illustrating a neural network training method provided in an embodiment of this application. The neural network training method provided in this embodiment may include:

[0136] 1001. Obtain training samples and corresponding second expectation information. The training samples include at least images of the vehicle's surrounding environment. The second expectation information includes correct images of the vehicle's surrounding environment from a top-down perspective. The second expectation information also indicates the correct category of objects in the correct images from a top-down perspective.

[0137] In this embodiment of the application, before training the second neural network, the training device needs to obtain at least one training sample from the training dataset. Each training sample includes at least one set of images, and the aforementioned set of images includes at least one image of the vehicle's surrounding environment. Optionally, each training sample may also include at least one point cloud data of the vehicle's surrounding environment. The specific images included in "at least one image of the vehicle's surrounding environment" and the specific point cloud data included in "at least one point cloud data of the vehicle's surrounding environment" can be understood in conjunction with the description in the embodiment corresponding to Figure 3 above, and will not be elaborated here.

[0138] The second expected information corresponding to the training samples can also be called the second label corresponding to the training samples, or the second ground truth corresponding to the training samples. Optionally, the second expected information corresponding to the training samples is obtained based on a standard-refinement map and a high-refinement map corresponding to the vehicle's surrounding environment. For example, the second expected information corresponding to the training samples includes the presence of roads and buildings in the correct image of the vehicle's surrounding environment from a top-down view; wherein, the image of the road portion in the correct image of the vehicle's surrounding environment from a top-down view can be obtained based on a high-refinement map, and the image of the building portion in the correct image of the vehicle's surrounding environment from a top-down view can be obtained based on a standard-refinement map; the correct category of objects on the road in the vehicle's surrounding environment can be obtained based on the high-refinement map corresponding to the vehicle's surrounding environment, and the correct category of buildings in the vehicle's surrounding environment can be obtained based on the standard-refinement map corresponding to the vehicle's surrounding environment.

[0139] In this embodiment, since the high-precision map carries detailed road information and the standard-precision map carries building information in the environment, and the second expected information includes a correct image of the vehicle's surrounding environment from a top-down perspective, and the second expected information also indicates the correct category of objects in the correct image from a top-down perspective, using both the high-precision map and the standard-precision map to generate the second expected information is beneficial because the second expected information contains both high-quality road information and building information in the environment. That is, the second expected information can more comprehensively and accurately reflect the vehicle's surrounding environment. Using the second expected information as a guide to train the second neural network is beneficial because the trained second neural network has a better understanding of the vehicle's surrounding environment. That is, it is beneficial because the first feature information extracted by the trained second neural network covers richer information, which in turn is beneficial because the first neural network can generate more accurate prediction information with the help of the first feature information and the second feature information.

[0140] Since subsequent steps 1003 and 1004 are optional, if steps 1003 and 1004 are executed, each training sample may also include a first location, and the training device may then obtain map information within a first preset range of the first location based on the first location; or, each training sample may also include map information within a first preset range of the first location. The meanings of "first location" and "map information within a first preset range of the first location" can be found in the descriptions in the embodiments corresponding to Figure 3 above, and will not be repeated here.

[0141] In step 1001, the training device also needs to acquire the first expected information corresponding to the training sample. The first expected information is used to determine the correct actual position of the vehicle. The first expected information corresponding to the training sample can also be called the first label corresponding to the training sample, or the first ground truth corresponding to the training sample.

[0142] For example, in one case, the first expected information indicates the offset between the vehicle's correct actual position and the first position; in another case, the first expected information indicates the vehicle's correct actual position. The meaning of "first expected information" is similar to that of "first predicted information," except that "predicted" in the description of the first predicted information is replaced with "correct." That is, the first predicted information is obtained by prediction through a neural network, while the first expected information is correct information.

[0143] Optionally, the correct actual position of the vehicle can be the position of the image of the vehicle's surrounding environment included in the training sample. That is, the training device can take the acquisition position of the image of the vehicle's surrounding environment as the correct actual position of the vehicle. The training device can also randomly generate first expected information, that is, randomly generate the offset between the correct actual position of the vehicle and the first position. Then, based on the correct actual position of the vehicle and the randomly generated first expected information, the first position is determined, thereby enabling the acquisition of map information within a first preset range of the first position.

[0144] 1002. Input the image of the vehicle's surrounding environment into the second neural network, and extract features from the image of the vehicle's surrounding environment through the second neural network to obtain the first feature information.

[0145] In this embodiment of the application, the specific implementation of step 1002 by the training device and the specific meaning of the terms in step 1002 can be found in the description of step 303 in the embodiment corresponding to Figure 3 above, and will not be repeated here.

[0146] 1003. Input the first feature information into the third neural network to obtain the second prediction information generated by the third neural network. The second prediction information includes the predicted image of the vehicle's surrounding environment from a top-down perspective. The second prediction information also indicates the predicted category of the object in the predicted image from a top-down perspective.

[0147] In this embodiment of the application, after the training device generates the first feature information through the second neural network, it can also input the first feature information into the third neural network to obtain the second prediction information generated by the third neural network.

[0148] For example, the third neural network can also be understood as a decoder. For instance, the third neural network can be specifically represented as an MLP, a convolutional neural network + MLP, or other network structures, which can be determined in combination with the actual application environment. This application does not limit the specific implementation.

[0149] The second prediction information can be understood as the semantic information of objects in the vehicle's surrounding environment from a top-down perspective. For example, the second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down perspective, and the second prediction information also indicates the predicted category of objects in the predicted image from a top-down perspective. For example, the second prediction information can be represented as a predicted image carrying annotation information, which is used to indicate the predicted category of objects in the predicted image from a top-down perspective.

[0150] 1004. Extract features from the map information within the first preset range of the first location to obtain second feature information. The first location represents the vehicle's location obtained by the vehicle positioning system, and the map information within the first preset range of the first location comes from a standard and refined map.

[0151] 1005. Based on the first feature information and the second feature information, a first prediction information is generated through a first neural network. The first prediction information is used to determine the actual position of the predicted vehicle.

[0152] In this embodiment of the application, steps 1004 and 1005 are optional steps. If steps 1004 and 1005 are executed, for example, step 1004 may include: the training device extracts features from the map information within a first preset range of the first location through a fourth neural network to obtain second feature information. It should be noted that the specific implementation of steps 1004 and 1005 by the training device and the specific meaning of the terms in steps 1004 and 1005 can be referred to the description of steps 304 and 305 in the embodiment corresponding to Figure 3 above, and will not be repeated here.

[0153] 1006. The second neural network is trained using the first loss function to obtain the trained second neural network. The first loss function indicates the similarity between the second predicted information and the second expected information.

[0154] In this embodiment, steps 1004 and 1005 are optional. If steps 1004 and 1005 are not executed, then in step 1006, the training device can use a first loss function to train the second neural network (optionally, also including a third neural network) until a first convergence condition is met, resulting in the trained second neural network (optionally, also including a trained third neural network). The first loss function indicates the similarity between the second predicted information and the second expected information. The goal of training using the second loss function includes improving the similarity between the second predicted information generated by the second neural network and the second expected information. The first convergence condition may include satisfying the convergence condition of the first loss function, and / or, the number of iterative training iterations reaching a preset number.

[0155] For example, during each training process of the second neural network (optionally, also including a third neural network), the training device can generate the function value of the first loss function based on the second prediction information and the second expectation information, and then update the weight parameters of the second neural network (optionally, also including a third neural network) using the backpropagation algorithm based on the function value of the first loss function, so as to realize one training of the second neural network (optionally, also including a third neural network).

[0156] For example, during the training of the second neural network, the weight vector of each layer can be updated based on the difference between the predicted value obtained from the second neural network and the desired expected value (of course, there is usually an initialization process before the first update, that is, pre-configuring the parameters of each layer in the second neural network). For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict lower, and this adjustment is continued until the second neural network can generate the truly desired expected value or a value very close to the truly desired expected value. Therefore, it is necessary to predefine "how to compare the difference between the predicted value and the expected value," which is the loss function or objective function. These are important equations used to measure the difference between the predicted value and the expected value. Taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. The training of the second neural network then becomes the process of minimizing this loss as much as possible.

[0157] Neural networks can employ backpropagation algorithms to refine the initial parameters during training, thereby minimizing reconstruction error. Specifically, forward propagation of the input signal to the output generates error loss. This error loss information is then used to update the initial neural network parameters, leading to convergence of the error loss. The backpropagation algorithm is an error-loss-driven process aimed at obtaining optimal neural network parameters.

[0158] To more intuitively understand this solution, please refer to Figure 11. Figure 11 is a schematic diagram of a neural network training method provided in an embodiment of this application. As shown in Figure 11, after acquiring training samples, the training device inputs images of the vehicle's surrounding environment, including the training samples, into a second neural network. The second neural network extracts features from the images of the vehicle's surrounding environment to obtain first feature information generated by the second neural network. The first feature information is then input into a third neural network to obtain second prediction information generated by the third neural network. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down perspective, and also indicates the predicted category of objects in the predicted image from the top-down perspective. The training device generates the function value of a first loss function based on the second prediction information and the second expected information. The second neural network and the third neural network are trained based on the function value of the first loss function. It should be understood that the example in Figure 11 is only for the convenience of understanding this solution and is not intended to limit this solution.

[0159] In this embodiment, during the training phase of the second neural network, the feature information generated by the second neural network is input into the third neural network. The third neural network generates second prediction information, which indicates the predicted image of the vehicle's surrounding environment from a top-down perspective and the predicted categories of objects in the image. Then, the second neural network is trained using a first loss function, which indicates the similarity between the second prediction information and the second expected information. The second expected information includes the correct image of the vehicle's surrounding environment and the correct categories of objects in the image. That is, based on the feature information of the vehicle's surrounding environment generated by the second neural network from a top-down perspective, semantic segmentation of the vehicle's surrounding environment is performed. The correct semantic segmentation result is used as supervision to train the second neural network. Since the better the second feature information collected by the second neural network, the simpler the semantic segmentation of the third neural network is, and the easier it is to generate better quality semantic segmentation results, the higher the similarity between the second prediction information and the second expected information. This training method is beneficial to improving the second neural network's ability to better understand the vehicle's surrounding environment, which also helps to ensure that the first feature information extracted by the trained second neural network contains richer information.

[0160] Optionally, if steps 1004 and 1005 are performed, step 1006 includes: the training device using a first loss function and a second loss function to train a first neural network and a second neural network (optionally, also including a third neural network and a fourth neural network) until a second convergence condition is met, resulting in a trained first neural network and a trained second neural network (optionally, also including a trained third neural network and a trained fourth neural network). The second loss function indicates the similarity between the first predicted information and the first expected information, where the first expected information is used to determine the correct actual position of the vehicle. The second convergence condition may include satisfying the convergence conditions of the first and second loss functions, and / or, the number of iterative training iterations reaching a preset number.

[0161] For example, in each training process of the first neural network and the second neural network (optionally, also including the third neural network and the fourth neural network), the training device can generate the function value of the first loss function based on the second prediction information and the second expectation information, generate the function value of the second loss function based on the first prediction information and the first expectation information, and then update the weight parameters of the first neural network and the second neural network (optionally, also including the third neural network and the fourth neural network) using the backpropagation algorithm based on the function values ​​of the first loss function and the second loss function, so as to realize one training of the first neural network and the second neural network (optionally, also including the third neural network and the fourth neural network).

[0162] Optionally, if steps 1004 and 1005 are performed, the training device can also obtain sub-feature information from the second feature information. The second feature information includes feature information of map information within a first preset range of the first location. The map information within the first preset range of the first location includes map information within a second preset range of the vehicle's actual location. The sub-feature information includes feature information of map information within a second preset range of the actual location. The meaning of "second preset range" is similar to that of "first preset range," except that "second preset range" is smaller than "first preset range." For example, the first preset range is a square area with a side length of 256 meters centered on the first location, while the second preset range is a square area with a side length of 128 meters centered on the actual location.

[0163] In step 1006, the training device uses a first loss function, a second loss function, and a third loss function to train the first neural network and the second neural network (optionally, also including a third neural network and a fourth neural network) until a third convergence condition is met, resulting in the trained first neural network and the trained second neural network (optionally, also including a trained third neural network and a trained fourth neural network). The third loss function indicates the similarity between the first feature information and the sub-feature information; the purpose of training using the third loss function is to improve the similarity between the first feature information and the sub-feature information. The third convergence condition may include satisfying the convergence conditions of the first loss function, the second loss function, and the third loss function, and / or, the number of iterative training iterations reaching a preset number.

[0164] In this embodiment, since the image of the vehicle's surrounding environment can reflect the vehicle's actual position, theoretically, the similarity between the image of the vehicle's surrounding environment and the map information within a preset range of the vehicle's actual position is relatively high. That is, the similarity between the feature information of the image of the vehicle's surrounding environment and the feature information of the map information within a preset range of the vehicle's actual position should also be relatively high. The training device introduces a third loss function during the training phase of the second neural network. The third loss function represents the similarity between the first feature information (i.e., the feature information of the image of the vehicle's surrounding environment) and the sub-feature information (i.e., the feature information within a second preset range of the vehicle's actual position). This supervises the training of the second neural network in more dimensions, which helps the trained second neural network to learn better feature extraction capabilities. In other words, it helps the second neural network to extract more accurate feature information from the image of the vehicle's surrounding environment, thereby facilitating more accurate vehicle positioning.

[0165] Based on the embodiments corresponding to Figures 1 to 11, in order to better implement the above-described solutions of this application, related equipment for implementing the above-described solutions is also provided below. Specifically, referring to Figure 12, Figure 12 is a structural schematic diagram of a vehicle positioning device provided in an embodiment of this application. The vehicle positioning device 1200 includes: an acquisition module 1201, used to acquire images of the vehicle's surrounding environment and map information within a preset range of a first location, the first location being the vehicle's location obtained based on a vehicle positioning system; a feature extraction module 1202, used to extract features from the images of the vehicle's surrounding environment to obtain first feature information; the feature extraction module 1202 is also used to extract features from the map information within the preset range of the first location to obtain second feature information, the map information within the preset range of the first location originating from a high-precision map; and a generation module 1203, used to generate first prediction information through a first neural network based on the first feature information and the second feature information, the first prediction information being used to determine the actual location of the vehicle.

[0166] Optionally, the first neural network includes a Transformer neural network module and a multilayer perceptron (MLP). The generation module 1203 is specifically used to: fuse the first feature information and the second feature information to obtain fused feature information; input the fused feature information into the first neural network to obtain the first prediction information generated by the first neural network, wherein the first prediction information indicates the offset between the actual position of the vehicle and the first position.

[0167] Optionally, the images of the vehicle's surrounding environment include at least two images of the vehicle's surrounding environment, which include images acquired by the vehicle in at least two of the following directions: left front, front, right front, left rear, rear, or right rear.

[0168] Optionally, the first feature information includes feature information of the vehicle's surrounding environment from a top-down perspective. The feature extraction module 1202 is specifically used to extract features from the image of the vehicle's surrounding environment through a second neural network. During the training of the second neural network using a first loss function, the feature information generated by the second neural network is also input into a third neural network. The third neural network is used to generate second prediction information. The first loss function indicates the similarity between the second prediction information and the second expected information. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down perspective, and the second prediction information also indicates the predicted category of objects in the predicted image from a top-down perspective. The second expected information includes a correct image of the vehicle's surrounding environment from a top-down perspective, and the second expected information also indicates the correct category of objects in the correct image from a top-down perspective.

[0169] Optionally, the vehicle positioning device 1200 further includes: a rasterization module 1204, used to rasterize the map information within a preset range of the first location in the first format to obtain the map information within a preset range of the first location in the second format, wherein the first format is a text format or a binary format and the second format is an image format; and a feature extraction module, specifically used to extract features from the map information within a preset range of the first location in the second format.

[0170] It should be noted that the information interaction and execution process between the modules / units in the vehicle positioning device 1200 are based on the same concept as the various method embodiments corresponding to Figures 1 to 11 in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0171] Referring to Figure 13, which is a schematic diagram of a neural network training device provided in an embodiment of this application, the neural network training device 1300 includes: an input module 1301, used to input an image of the vehicle's surrounding environment into a second neural network, and extract features from the image of the vehicle's surrounding environment through the second neural network to obtain first feature information; a generation module 1302, used to input the first feature information into a third neural network to obtain second prediction information generated by the third neural network, the second prediction information including a predicted image of the vehicle's surrounding environment from a top-down perspective, and the second prediction information also indicating the predicted category of objects in the predicted image from a top-down perspective; and a training module 1303, used to train the second neural network using a first loss function to obtain a trained second neural network, the first loss function indicating the similarity between the second prediction information and the second expected information, the second expected information including a correct image of the vehicle's surrounding environment from a top-down perspective, and the second expected information also indicating the correct category of objects in the correct image from a top-down perspective.

[0172] Optionally, the second desired information is obtained based on a standard-precision map and a high-precision map corresponding to the vehicle's surrounding environment.

[0173] Optionally, the training device 1300 for the neural network further includes: a feature extraction module 1304, used to extract features from map information within a preset range of the first location to obtain second feature information, wherein the first location represents the location of the vehicle obtained by the vehicle-based positioning system, and the map information within the preset range of the first location is derived from a high-precision map; a generation module 1302, further used to generate first prediction information through a first neural network based on the first feature information and the second feature information, wherein the first prediction information is used to determine the predicted actual location of the vehicle; and a training module 1304, specifically used to train the first neural network and the second neural network using a first loss function and a second loss function to obtain the trained first neural network and the trained second neural network, wherein the second loss function indicates the similarity between the first prediction information and the first expected information, and the first expected information is used to determine the correct actual location of the vehicle.

[0174] It should be noted that the information interaction and execution process between the modules / units in the neural network training device 1300 are based on the same concept as the various method embodiments corresponding to Figures 1 to 11 in this application. For details, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.

[0175] The following describes a device provided in an embodiment of this application. When this device is specifically implemented as an execution device, please refer to Figure 14. Figure 14 is a structural schematic diagram of the device provided in an embodiment of this application. Specifically, device 1400 includes: receiver 1401, transmitter 1402, processor 1403, and memory 1404 (wherein the number of processors 1403 in device 1400 can be one or more; Figure 14 shows one processor as an example). The processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of this application, receiver 1401, transmitter 1402, processor 1403, and memory 1404 can be connected via a bus or other means.

[0176] Memory 1404 may include read-only memory and random access memory, and provides instructions and data to processor 1403. A portion of memory 1404 may also include non-volatile random access memory (NVRAM). Memory 1404 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.

[0177] Processor 1403 controls the operation of the device. In specific applications, the various components of the device are coupled together through a bus system, which may include not only the data bus, but also power buses, control buses, and status signal buses. However, for clarity, all buses in the diagram are referred to as the bus system.

[0178] The methods disclosed in the embodiments of this application can be applied to or implemented by the processor 1403. The processor 1403 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuits in the hardware of the processor 1403 or by instructions in software form. The processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1403 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can reside in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 1404. Processor 1403 reads the information in memory 1404 and, in conjunction with its hardware, completes the steps of the above method.

[0179] Receiver 1401 can be used to receive input digital or character information, and to generate signal inputs related to device settings and function control. Transmitter 1402 can be used to output digital or character information through the first interface; transmitter 1402 can also be used to send instructions to the disk group through the first interface to modify data in the disk group; transmitter 1402 may also include a display device such as a display screen.

[0180] In this embodiment, processor 1403 is used to execute the vehicle execution method in the embodiments corresponding to Figures 1 to 11. It should be noted that the specific manner in which the application processor 14031 in processor 1403 executes the aforementioned steps is based on the same concept as the method embodiments corresponding to Figures 1 to 11 in this application, and the resulting technical effects are the same as those in the method embodiments corresponding to Figures 1 to 11 in this application. For details, please refer to the descriptions in the method embodiments shown above in this application, which will not be repeated here.

[0181] In the case where the device specifically manifests as a second device, please refer to Figure 15. Figure 15 is a schematic diagram of another structure of the device provided in the embodiment of this application. Specifically, device 1500 is implemented by one or more servers. Device 1500 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1522 (e.g., one or more processors) and memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) for storing application programs 1542 or data 1544. The memory 1532 and storage media 1530 can be temporary or persistent storage. The program stored in storage media 1530 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the device. Furthermore, the CPU 1522 may be configured to communicate with storage media 1530 and execute a series of instruction operations in storage media 1530 on device 1500.

[0182] Device 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0183] In this embodiment, the central processing unit 1522 is used to execute the method of the training device in the embodiments corresponding to Figures 1 to 11. It should be noted that the specific manner in which the central processing unit 1522 executes the above steps is based on the same concept as the method embodiments corresponding to Figures 1 to 11 in this application, and the resulting technical effects are the same as those in the method embodiments corresponding to Figures 1 to 11 in this application. For details, please refer to the description in the aforementioned method embodiments of this application, which will not be repeated here.

[0184] This application also provides a vehicle, as shown in Figure 16. Figure 16 is a structural schematic diagram of a vehicle provided in this application embodiment. The vehicle 100 is configured for fully or partially automated driving mode. For example, the vehicle 100 can control itself while in automated driving mode, and can determine the current state of the vehicle and its surrounding environment through human operation, determine the possible behavior of at least one other vehicle in the surrounding environment, and determine the confidence level corresponding to the probability of other vehicles performing possible behaviors, and control the vehicle 100 based on the determined information. When the vehicle 100 is in automated driving mode, the vehicle 100 can also be set to operate without human interaction.

[0185] Vehicle 100 may include various subsystems, such as a mobility system 102, a sensor system 104, a control system 106, one or more peripheral devices 108, a power supply 110, a computer system 112, and a user interface 116. Optionally, vehicle 100 may include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of vehicle 100 may be interconnected via wired or wireless means.

[0186] The mobility system 102 may include components that provide powered motion to the vehicle 100. In one embodiment, the mobility system 102 may include an engine 118, an energy source 119, a transmission 120, and wheels / tires 121.

[0187] Engine 118 can be an internal combustion engine, an electric motor, an air-compressed engine, or other combinations of engines, such as a hybrid engine consisting of a gasoline engine and an electric motor, or a hybrid engine consisting of an internal combustion engine and an air-compressed engine. Engine 118 converts energy source 119 into mechanical energy. Examples of energy source 119 include gasoline, diesel, other petroleum-based fuels, propane, other compressed gas-based fuels, ethanol, solar panels, batteries, and other sources of electricity. Energy source 119 can also provide energy to other systems of vehicle 100. Transmission 120 transmits mechanical power from engine 118 to wheels 121. Transmission 120 may include a gearbox, a differential, and a drive shaft. In one embodiment, transmission 120 may also include other components, such as a clutch. The drive shaft may include one or more axles that can be coupled to one or more wheels 121.

[0188] Sensor system 104 may include several sensors for sensing information about the environment surrounding vehicle 100. For example, sensor system 104 may include a positioning system 122 (which may be a GPS system, a BeiDou system, or another positioning system), an inertial measurement unit (IMU) 124, a radar 126, a laser rangefinder 128, and a camera 130. Sensor system 104 may also include sensors for the internal systems of the monitored vehicle 100 (e.g., an in-vehicle air quality monitor, fuel gauge, oil temperature gauge, etc.). Sensing data from one or more of these sensors can be used to detect objects and their corresponding characteristics (position, shape, orientation, speed, etc.). This detection and identification is a key function for the safe operation of the autonomous vehicle 100.

[0189] The positioning system 122 can be used to estimate the geographical location of the vehicle 100. An IMU 124 is used to sense changes in the position and orientation of the vehicle 100 based on inertial acceleration. In one embodiment, the IMU 124 can be a combination of an accelerometer and a gyroscope. A radar 126 can use radio signals to sense objects in the surrounding environment of the vehicle 100, specifically millimeter-wave radar or lidar. In some embodiments, in addition to sensing objects, the radar 126 can also be used to sense the speed and / or direction of travel of objects. A laser rangefinder 128 can use lasers to sense objects in the environment in which the vehicle 100 is located. In some embodiments, the laser rangefinder 128 may include one or more laser sources, a laser scanner, and one or more detectors, as well as other system components. A camera 130 can be used to capture multiple images of the surrounding environment of the vehicle 100. The camera 130 can be a still camera or a video camera.

[0190] The control system 106 controls the operation of the vehicle 100 and its components. The control system 106 may include various components, including a steering system 132, a throttle 134, a braking unit 136, a computer vision system 140, a trajectory control system 142, and an obstacle avoidance system 144.

[0191] The steering system 132 is operable to adjust the forward direction of the vehicle 100. For example, in one embodiment, it may be a steering wheel system. The throttle 134 controls the operating speed of the engine 118 and thus the speed of the vehicle 100. The braking unit 136 controls the deceleration of the vehicle 100. The braking unit 136 may use friction to slow down the wheels 121. In other embodiments, the braking unit 136 may convert the kinetic energy of the wheels 121 into electrical current. The braking unit 136 may also take other forms to slow down the rotational speed of the wheels 121 to control the speed of the vehicle 100. The computer vision system 140 is operable to process and analyze images captured by the camera 130 to identify objects and / or features in the environment surrounding the vehicle 100. The objects and / or features may include traffic signals, road boundaries, and obstacles. The computer vision system 140 may use object recognition algorithms, Structure from Motion (SFM) algorithms, video tracking, and other computer vision techniques. In some embodiments, the computer vision system 140 may be used to map the environment, track objects, estimate the speed of objects, etc. The route control system 142 is used to determine the driving route and speed of the vehicle 100. In some embodiments, the route control system 142 may include a lateral planning module 1421 and a longitudinal planning module 1422, which are respectively used to combine data from the obstacle avoidance system 144, GPS 122, and one or more predetermined maps to determine the driving route and speed for the vehicle 100. The obstacle avoidance system 144 is used to identify, evaluate, and avoid or otherwise traverse obstacles in the environment of the vehicle 100, which may specifically be physical obstacles and virtual moving bodies that may collide with the vehicle 100. In one example, the control system 106 may add or alternatively include components other than those shown and described. Alternatively, some of the components shown above may be reduced.

[0192] Vehicle 100 interacts with external sensors, other vehicles, other computer systems, or users via peripheral device 108. Peripheral device 108 may include wireless communication system 146, on-board computer 148, microphone 150, and / or speaker 152. In some embodiments, peripheral device 108 provides a means for a user of vehicle 100 to interact with user interface 116. For example, on-board computer 148 may provide information to a user of vehicle 100. User interface 116 may also operate on-board computer 148 to receive user input. On-board computer 148 may be operated via a touchscreen. In other cases, peripheral device 108 may provide a means for vehicle 100 to communicate with other devices located within the vehicle. For example, microphone 150 may receive audio (e.g., voice commands or other audio input) from a user of vehicle 100. Similarly, speaker 152 may output audio to a user of vehicle 100. Wireless communication system 146 may communicate wirelessly with one or more devices, either directly or via a communication network. For example, the wireless communication system 146 may use 3G cellular communication, such as CDMA, EVDO, GSM / GPRS, or 4G cellular communication, such as LTE, or 5G cellular communication. The wireless communication system 146 may utilize a wireless local area network (WLAN) for communication. In some embodiments, the wireless communication system 146 may utilize an infrared link, Bluetooth, or ZigBee to communicate directly with the device. Other wireless protocols, such as various vehicle communication systems, may also be used. For example, the wireless communication system 146 may include one or more dedicated short-range communications (DSRC) devices that can enable public and / or private data communication between the vehicle and / or a roadside station.

[0193] Power source 110 can provide power to various components of vehicle 100. In one embodiment, power source 110 can be a rechargeable lithium-ion or lead-acid battery. One or more such battery packs can be configured to provide power to various components of vehicle 100. In some embodiments, power source 110 and energy source 119 can be implemented together, as is the case in some fully electric vehicles.

[0194] Some or all of the functions of vehicle 100 are controlled by computer system 112. Computer system 112 may include at least one processor 113, which executes instructions 115 stored in a non-transitory computer-readable medium such as memory 114. Computer system 112 may also be multiple computing devices that control individual components or subsystems of vehicle 100 in a distributed manner. Processor 113 may be any conventional processor, such as a commercially available central processing unit (CPU). Alternatively, processor 113 may be a dedicated device such as an application-specific integrated circuit (ASIC) or other hardware-based processor. Although FIG16 functionally illustrates the processor, memory, and other components of computer system 112 in the same block, those skilled in the art will understand that the processor or memory may actually include multiple processors or memories not stored in the same physical housing. For example, memory 114 may be a hard disk drive or other storage medium located in a housing different from that of computer system 112. Therefore, references to processor 113 or memory 114 will be understood to include references to a collection of processors or memories that may or may not operate in parallel. Unlike using a single processor to perform the steps described herein, some components, such as steering and deceleration components, can each have their own processor that performs only calculations related to the component's specific function.

[0195] In all the aspects described herein, processor 113 may be located remotely from vehicle 100 and may communicate wirelessly with vehicle 100. In other aspects, some of the processes described herein are executed on processor 113 located within vehicle 100, while others are executed by remote processor 113, including taking the necessary steps to perform a single operation.

[0196] In some embodiments, memory 114 may contain instructions 115 (e.g., program logic) that can be executed by processor 113 to perform various functions of vehicle 100, including those described above. Memory 114 may also contain additional instructions, including instructions to send data to, receive data from, interact with, and / or control one or more of the mobility system 102, sensor system 104, control system 106, and peripheral devices 108. In addition to instructions 115, memory 114 may also store data such as road maps, route information, vehicle position, direction, speed, and other such vehicle data, as well as other information. This information may be used by vehicle 100 and computer system 112 during operation of vehicle 100 in autonomous, semi-autonomous, and / or manual modes. A user interface 116 is provided to or receives information from a user of vehicle 100. Optionally, user interface 116 may include one or more input / output devices within the set of peripheral devices 108, such as wireless communication system 146, on-board computer 148, microphone 150, and speaker 152.

[0197] Computer system 112 can control the functions of vehicle 100 based on input received from various subsystems (e.g., driving system 102, sensor system 104, and control system 106) and from user interface 116. For example, computer system 112 can utilize input from control system 106 to control steering system 132 to avoid obstacles detected by sensor system 104 and obstacle avoidance system 144. In some embodiments, computer system 112 is operable to provide control over many aspects of vehicle 100 and its subsystems.

[0198] Alternatively, one or more of these components may be installed separately from or associated with vehicle 100. For example, memory 114 may exist partially or completely separately from vehicle 100. The components may be communicatively coupled together in a wired and / or wireless manner.

[0199] Optionally, the above components are merely examples. In practical applications, components in each of the above modules may be added or removed according to actual needs. Figure 16 should not be construed as a limitation on the embodiments of this application. A vehicle traveling on a road, such as vehicle 100 above, can identify objects in its surrounding environment to determine an adjustment to its current speed. These objects can be other vehicles, traffic control equipment, or other types of objects. In some examples, each identified object can be considered independently, and based on the object's individual characteristics, such as its current speed, acceleration, and distance from the vehicle, the speed adjustment to be made by the vehicle can be determined.

[0200] Optionally, vehicle 100 or computing devices associated with vehicle 100, such as computer system 112, computer vision system 140, and memory 114 as shown in Figure 16, can predict the behavior of the identified objects based on the characteristics of the identified objects and the state of the surrounding environment (e.g., traffic, rain, ice on the road, etc.). Optionally, each identified object depends on the behavior of each other, so all identified objects can also be considered together to predict the behavior of a single identified object. Vehicle 100 can adjust its speed based on the predicted behavior of the identified objects. In other words, vehicle 100 can determine what steady state the vehicle will need to adjust to (e.g., accelerate, decelerate, or stop) based on the predicted behavior of the objects. In this process, other factors can also be considered in determining the speed of vehicle 100, such as the lateral position of vehicle 100 in the road, the curvature of the road, the proximity of static and dynamic objects, etc. In addition to providing instructions to adjust the speed of the vehicle, the computing device can also provide instructions to modify the steering angle of the vehicle 100 so that the vehicle 100 follows a given trajectory and / or maintains a safe lateral and longitudinal distance from objects near the vehicle 100 (e.g., cars in adjacent lanes on the road).

[0201] In this embodiment, the processor 113 in the vehicle 100 is used to execute the methods performed by the vehicle in the embodiments corresponding to Figures 1 to 11. It should be noted that the specific manner in which the processor 113 executes the aforementioned steps is based on the same concept as the method embodiments corresponding to Figures 1 to 11 in this application, and the resulting technical effects are the same as those in the method embodiments corresponding to Figures 1 to 11 in this application. For details, please refer to the descriptions in the method embodiments shown above in this application; they will not be repeated here.

[0202] This application also provides a computer-readable storage medium storing a program that, when run on a computer, causes the computer to perform the steps performed by the vehicle in the methods described in the embodiments shown in Figures 1 to 11, or causes the computer to perform the steps performed by the training device in the methods described in the embodiments shown in Figures 1 to 11.

[0203] This application also provides a computer program product, which includes a program that, when run on a computer, causes the computer to perform the steps performed by the vehicle in the methods described in the embodiments shown in Figures 1 to 11, or causes the computer to perform the steps performed by the training device in the methods described in the embodiments shown in Figures 1 to 11.

[0204] This application also provides a circuit system including a processing circuit configured to perform the method described in the embodiments shown in Figures 1 to 11 above.

[0205] The execution device, training device, or vehicle positioning device provided in this application embodiment can specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, pins, or circuits. The processing unit can execute computer execution instructions stored in the storage unit to cause the chip to execute the methods described in the embodiments shown in Figures 1 to 11. Optionally, the storage unit is a storage unit within the chip, such as a register or cache. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, such as random access memory (RAM).

[0206] The processor mentioned above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits used to control the execution of a program in the first aspect of the method.

[0207] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the device embodiment drawings provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.

[0208] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CLUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0209] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, in the form of a computer program product.

[0210] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

Claims

1. A method for locating a vehicle, characterized in that, The method includes: Acquire images of the vehicle's surrounding environment and map information within a preset range of a first location, where the first location is the vehicle's position obtained based on the vehicle's positioning system; Feature extraction is performed on the image of the environment surrounding the vehicle to obtain first feature information; Feature extraction is performed on the map information within a preset range of the first location to obtain second feature information, wherein the map information within the preset range of the first location is derived from a high-precision map. Based on the first feature information and the second feature information, a first prediction information is generated through a first neural network, and the first prediction information is used to determine the actual position of the vehicle.

2. The method according to claim 1, characterized in that, The first neural network includes a Transformer neural network module and a Multilayer Perceptron (MLP). The step of generating first prediction information based on the first feature information and the second feature information through the first neural network includes: The first feature information and the second feature information are fused to obtain the fused feature information; The fused feature information is input into the first neural network to obtain the first prediction information generated by the first neural network. The first prediction information indicates the offset between the actual position of the vehicle and the first position.

3. The method according to claim 1 or 2, characterized in that, The images of the vehicle's surrounding environment include at least two images of the vehicle's surrounding environment, which include images captured by the vehicle in at least two of the following directions: left front, front, right front, left rear, rear, or right rear.

4. The method according to claim 1 or 2, characterized in that, The first feature information includes feature information of the vehicle's surrounding environment from a top-down perspective, and the feature extraction of the image of the vehicle's surrounding environment includes: The second neural network extracts features from the image of the vehicle's surrounding environment. During the training of the second neural network using the first loss function, the feature information generated by the second neural network is also input into the third neural network. The third neural network is used to generate second prediction information. The first loss function indicates the similarity between the second prediction information and the second expected information. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down perspective, and the second prediction information also indicates the predicted category of objects in the predicted image from the top-down perspective. The second desired information includes a correct image of the vehicle's surroundings from a top-down perspective, and the second desired information also indicates the correct category of objects in the correct image from the top-down perspective.

5. The method according to claim 1 or 2, characterized in that, The method further includes: The map information within a preset range of the first location in the first format is rasterized to obtain the map information within a preset range of the first location in the second format. The first format is a text format or a binary format, and the second format is an image format. The step of extracting features from map information within a preset range of the first location includes: extracting features from map information within a preset range of the first location in the second format.

6. A method for training a neural network, characterized in that, The method includes: An image of the vehicle's surrounding environment is input into a second neural network, and the second neural network extracts features from the image of the vehicle's surrounding environment to obtain first feature information. The first feature information is input into the third neural network to obtain the second prediction information generated by the third neural network. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down perspective. The second prediction information also indicates the predicted category of objects in the predicted image from the top-down perspective. The second neural network is trained using a first loss function to obtain a trained second neural network. The first loss function indicates the similarity between the second predicted information and the second expected information. The second expected information includes a correct image of the vehicle's surrounding environment from a top-down perspective. The second expected information also indicates the correct category of the object in the correct image from the top-down perspective.

7. The method according to claim 6, characterized in that, The second expected information is obtained based on a standard-precision map and a high-precision map corresponding to the environment surrounding the vehicle.

8. The method according to claim 6 or 7, characterized in that, The method further includes: Feature extraction is performed on the map information within a preset range of the first location to obtain second feature information. The first location represents the vehicle's location obtained by the vehicle's positioning system, and the map information within the preset range of the first location comes from a high-precision map. Based on the first feature information and the second feature information, a first prediction information is generated through a first neural network. The first prediction information is used to determine the predicted actual position of the vehicle. The step of training the second neural network using a first loss function to obtain the trained second neural network includes: The first neural network and the second neural network are trained using the first loss function and the second loss function to obtain the trained first neural network and the trained second neural network. The second loss function indicates the similarity between the first predicted information and the first expected information, and the first expected information is used to determine the correct actual location of the vehicle.

9. A vehicle positioning device, characterized in that, The device includes: The acquisition module is used to acquire images of the vehicle's surrounding environment and map information within a preset range of a first location, wherein the first location is the vehicle's position obtained based on the vehicle's positioning system. The feature extraction module is used to extract features from the image of the environment surrounding the vehicle to obtain first feature information; The feature extraction module is further configured to extract features from the map information within a preset range of the first location to obtain second feature information, wherein the map information within the preset range of the first location is derived from a high-precision map. The generation module is used to generate first prediction information based on the first feature information and the second feature information through a first neural network. The first prediction information is used to determine the actual position of the vehicle.

10. A training device for a neural network, characterized in that, The device includes: The input module is used to input an image of the vehicle's surrounding environment into a second neural network, and to extract features from the image of the vehicle's surrounding environment through the second neural network to obtain first feature information. The input module is further configured to input the first feature information into the third neural network to obtain the second prediction information generated by the third neural network. The second prediction information includes a predicted image of the vehicle's surrounding environment from a top-down perspective. The second prediction information also indicates the predicted category of objects in the predicted image from the top-down perspective. A training module is used to train the second neural network using a first loss function to obtain a trained second neural network. The first loss function indicates the similarity between the second predicted information and the second expected information. The second expected information includes a correct image of the vehicle's surrounding environment from a top-down perspective. The second expected information also indicates the correct category of the object in the correct image from the top-down perspective.

11. A device, characterized in that, The method includes a processor coupled to a memory storing program instructions, which, when executed by the processor, implement the method of any one of claims 1 to 8.

12. A vehicle, characterized in that, The method includes a processor coupled to a memory storing program instructions, which, when executed by the processor, implement the method of any one of claims 1 to 5.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 8.

14. A computer program product, characterized in that, The computer program product includes a program that, when run on a computer, causes the computer to perform the method as described in any one of claims 1 to 8.